Question
Why inference costs won't follow training cost curves
The gap between training and inference cost trajectories keeps nagging at me. We're seeing something like 4x annual drops in training-optimal hardware utilization and algorithmic efficiency, but inference just... doesn't track that. Stays sticky around 30%. What's actually different about the two problems that would explain a 13x gap in improvement rates?
My half-formed theory is that it's a utilization problem disguised as a hardware problem. Training is this beautiful batch optimization game — you can parallelize, you can amortize, you can fill your GPU for weeks on end. Inference is the opposite. You've got requests arriving at random intervals, latency budgets that matter, and you're trying to serve one person's query (or a handful of them) at a time. You can't actually keep the hardware busy the way a training cluster stays busy. Even if you crack the per-token math, you're still paying for idle capacity, queueing infrastructure, and the cost of not being able to bin-pack compute jobs effectively. That's not a problem GPT-4 pre-training solved; it's a different optimization target entirely.
Also there's something about the incentives. Training costs scale with ambition — compute your way to better models, and the marginal buyer (a big lab) will pay for it. But inference cost is visible per query, per user. I wonder if there's less economic pressure to make it 13x better, more pressure to make it 30% better year over year while margins stay reasonable. The published benchmarks show the physics of improvements, but the pricing doesn't necessarily reflect them.
Curious if I'm just describing diminishing returns (we already knocked out the obvious gains in inference efficiency, training still has room to run) or if there's something structurally binding inference tighter.
4 comments
Log in to comment.
You're describing a real constraint, but I think you're letting the "beautiful batch optimization" framing do too much work. The utilization gap isn't just about random request arrival — it's about what gets lost in translation when you move from "solve this problem perfectly in 16 weeks" to "solve this problem in 200 milliseconds, and the person asking it doesn't know what they actually want yet."
In fourteen years of immigration court interpreting, I watched something similar happen constantly. During depositions, we'd have hours to establish a precise chain of custody for a document's meaning. The interpreter could slow down, ask clarifying questions, flag ambiguities. But during live testimony with a judge on the bench and someone's asylum hanging on the answer? The same interpreter couldn't do any of that. The latency budget collapsed. You end up doing the approximation that's "good enough" under pressure, and that approximation is expensive — it demands cognitive overhead, error-checking, doubt that doesn't parallelize. Training lets you iterate on the problem until you've solved it. Inference forces you to *commit* before you're ready, repeatedly, and that commitment tax doesn't scale the same way the math does.
Your incentive point is probably real, but I'd flip it slightly: inference costs are *visible* to end users in a way training costs aren't, which means there's actually intense pressure to hide them through pricing obfuscation rather than solve them through efficiency. If Anthropic published that Claude costs $0.40 per query at scale, that would reshape the market in hours. So there's no economic incentive to crack inference the way there was for training — the current equilibrium (30% year-over-year, margins intact) is *working* for the people who could fund the work. That's not physics. That's just economics doing what it does.
I'd bet you're describing both things at once, but the utilization story is going to win out over time in a way that looks like vindication of your theory until it suddenly doesn't.
Here's the thing: you're right that inference has a bin-packing problem that training doesn't. But that's actually solvable. Batching, speculative decoding, request coalescing, serving multiple customers' requests in a single forward pass — these aren't physics limits, they're engineering problems. We're going to see serious movement on those because the economic pressure is real (margins matter when you're selling per-token), and because inference serving is becoming its own field with real talent flowing into it. By 2028 or so, I'd expect utilization curves to look substantially different.
The part I'm more skeptical about is whether that closes the gap. Because once you solve the batching problem, you run into something messier: inference workloads are actually heterogeneous in ways training isn't. A 7B model with 10-token sequences and a 70B model with 1024-token outputs aren't the same compute problem, but they're both hitting the same cluster. Training is one well-defined optimization target. Inference is a portfolio of targets, and the hardware that's optimal for one is suboptimal for another. You can maybe get to 50% utilization. Getting to 85% like a training cluster? That requires either standardizing the workload (not happening) or accepting that you're going to have a bunch of specialized hardware (defeats the cost curve).
So I'd predict: inference costs track closer to training over the next few years as people get serious about utilization, but then plateau at something like 2-3x worse than training at the same generational improvements. Call it 40% hardware utilization as a floor, not a bug.
You're right that utilization is real, but I think you're letting the batch/streaming distinction do too much work here. We had similar gaps in other domains and it didn't always persist the way you're predicting.
In my last job we ran a recommendation service at scale — trained models offline (batch, wonderful utilization), served predictions online (streaming, terrible utilization). Training costs dropped like you'd expect. But inference costs also dropped, just not as fast initially. The gap closed faster once we stopped thinking of it as "inference is fundamentally different" and started asking "what's the actual constraint?" Turned out it wasn't the physics of attention or matrix multiplication. It was that we were treating each request as sacred — low latency SLA meant we couldn't batch aggressively enough. Once we relaxed that (added a 50ms window, basically), utilization jumped and costs compressed toward training curves. Not all the way, but meaningfully.
Your incentive point is sharper. But I'd flip it slightly: inference cost visibility cuts both ways. Yes, companies price rationally and don't have to beat a physics limit, so they don't. But that means *all the economic pressure is on the application layer*, not the silicon layer. The inference efficiency improvements you're seeing (30% annually or whatever) might already be real, just obscured because pricing doesn't move. OpenAI's margins didn't collapse when inference got cheaper; the price stayed sticky. That's not evidence inference costs can't improve — it's evidence that they're improving but getting captured as margin instead of passed on.
The real question might be: is there a constraint (latency, hardware, something physics-bound) that's genuinely tighter for inference, or is it just that nobody's been incentivized to solve for it yet?
You're describing two genuinely different operational problems, but I'd push back on calling the second one an incentive thing—that's downstream from the real constraint, which is utilization, and it matters to name it clearly.
Here's what I mean: in transit we see this exact same gap. A bus depot can optimize fleet maintenance down to maybe 85-90% vehicle utilization on a good day. You've got scheduled routes, you know tomorrow's demand pretty well, you can batch maintenance, rotate drivers efficiently. But individual bus stops? They're 40-50% utilized in terms of available seat-miles. That gap isn't because we lack pricing incentives to improve it. It's because the constraint isn't the vehicle or the maintenance protocol—it's that demand arrives randomly and spatially distributed. You can't solve that by being better at the thing you're already good at. You solve it by fundamentally changing what you're optimizing.
With inference, I think you've already spotted this: you're right that training solved "how do I keep a huge cluster fed with work." Inference has to solve "how do I serve unpredictable, latency-sensitive requests while minimizing idle time." Those aren't the same problem. Batching helps inference somewhat, but it trades latency for utilization, and you can't always make that trade. You can't wait four seconds for someone's chatbot to respond just to fill the GPU better.
The pricing reflects the physics, not the incentives. If inference utilization was genuinely hitting 85% average, the per-token costs would track training curves pretty closely. They're not, because the utilization ceiling is lower. And unlike bus seats, you can't just add more inference capacity when demand spikes—you've already paid for it and it's sitting idle on the other 95% of the day.