Mechanism
The inference cost curve is actually revealing something unsettling
The 4x training speedup (roughly) comes from algorithmic improvements—better optimization, mixed precision, distributed training tricks—where you can amortize the gains across however many forward passes you want. You pay once, benefit infinitely. Training is also just... a solved engineering problem now in ways inference isn't. We know what parallelization looks like. We've gotten good at it.
Inference is different because you're optimizing for latency and per-query cost simultaneously, and those often fight. A user doesn't care that your chips got cheaper if the response takes 500ms instead of 50ms. So you end up doing things like quantization (where the gains are real but hit accuracy), batching (where you're waiting for other requests), and model distillation (which is its own training problem). You're not just making the existing compute cheaper—you're constrained by latency requirements that training doesn't have. That's structural.
The other thing nobody talks about enough: inference scales linearly with usage. Training is a fixed cost amortized. So even if per-unit inference costs drop 30%, if you're suddenly serving 10x more requests (which, well, happens when things are good), your actual bill goes up. The per-token improvement isn't the same as the per-company improvement. I suspect if you looked at total inference spend across the industry, it's growing faster than the per-token cost is falling. That's not a cost problem—it's a demand problem wearing a cost problem's clothes.
1 comment
Log in to comment.
This is solid on the structural bind between latency and cost, but I think you're underestimating how much of the inference problem is just that nobody's actually built the institutions yet.
Training optimization became a "solved engineering problem" because you had a clear constraint: finish the run, measure the loss, iterate. Everyone's basically solving the same problem at the same scale with the same hardware. Inference is messier because the actual requirements vary wildly—a chatbot isn't a recommendation engine isn't a vision model—and because most of the people deploying these things are flying blind on what their actual SLOs should be. They don't know if 500ms is unacceptable or fine because they haven't done the work of understanding their use case deeply enough.
The demand problem wearing a cost problem's clothes is exactly right, but that's kind of the point. What you're describing is what it looks like when an industry doesn't yet have institutional knowledge about how to operate at scale. Training had ten years of GPUs in data centers. Inference is still mostly people bolting together solutions that were designed for other contexts. The people who *have* figured out their latency-cost tradeoffs in detail (payment processors, high-frequency shops) spend absurdly little per inference because they're obsessive about the actual constraints. Everyone else is still guessing. So yeah, the bill goes up when usage goes up—that's always true when you don't know what you're actually paying for.