Question
Why inference costs aren't dropping like training costs
The standard story is that AI is getting cheaper across the board. Training costs fall 4x a year, so inference must be next. But the empirical picture is messier than that, and I think the structural differences matter more than people assume.
Training has this one magical property: you do it once, then amortize across millions of inference calls. So a 10x improvement in training efficiency *looks* like a 10x improvement in the per-inference cost of that model, even if the actual inference hardware or ops haven't changed much. Helion or Epoch's scaling laws show training FLOP requirements dropping fast. But that's partly a free lunch from better algorithms and scale—not necessarily from cheaper inference hardware. When people cite 4x/year improvements, they're often conflating two things: (1) training gets more efficient, so new models cost less to produce, and (2) running those models gets cheaper. Those aren't the same.
Inference hits different constraints. You need latency guarantees—can't batch forever on a consumer API. Memory bandwidth becomes the bottleneck, not raw compute. A single query to Claude or GPT-4 still needs to move model weights through memory, and memory bandwidth improvements lag FLOP improvements by years. You're also stuck with the installed base problem: training cost is capital spending that sinks, but inference infrastructure is operational. If you've got a data center full of H100s doing inference, you're not immediately retooling for cheaper hardware the way an organization might wait for the next generation of training chips.
The real question might be: is the gap closing at all? Or are we seeing improved training efficiency flatten against inference's actual hardware constraints?
4 comments
Log in to comment.
You're conflating two separate things here, and it's worth untangling because it changes what we should actually expect.
The training efficiency gains you're citing (Epoch's numbers, etc.) are mostly about algorithmic improvements and scale—better architectures, better data, longer training runs on the same hardware. That's real, but it's different from "inference gets cheaper." When a new model trains more efficiently, the *amortized per-inference cost* of that model drops, sure, but only if you're comparing it to running the old, less-efficient model. The actual hardware doing inference hasn't gotten proportionally better. You've just shifted to a smaller model or a more efficient architecture for the same task.
Inference hardware improvements are lagging because of exactly what you said: memory bandwidth is genuinely the bottleneck now, not multiply-accumulate throughput. An H100 can do enormous FLOPs, but you're bottlenecked moving weights around. We've known this for years—it's why folks have been pushing quantization, sparsity, and attention optimizations so hard. But those are *algorithmic workarounds* for a hardware constraint, not solutions to it. And yeah, the operational stickiness of deployed inference hardware means upgrade cycles are slow. I've watched teams run K80s for inference years after they stopped training anything on them, just because ripping out infrastructure is expensive and risky.
The gap-closing question is the right one, but I'd flip it: we're probably seeing training costs drop faster than inference costs for structural reasons that won't fix themselves just by waiting for better chips.
I've watched this exact thing happen in textiles and apparel. You'd get a new loom that cut production time by 40%, and everyone would say "great, unit costs drop 40%." But you still had to move finished goods through finishing, inspection, packing, shipping—all the downstream ops that the loom speedup didn't touch. So the *actual* cost reduction was maybe 12%, and it got smaller the more you optimized upstream. The constraint just moved, it didn't vanish.
What strikes me about the inference story is that people keep assuming the bottleneck is where it *was*, not where it *is*. If memory bandwidth is the real limit now, then throwing faster chips at training doesn't help you much on the API side. You need different hardware entirely—wider buses, faster memory—and that's a totally separate, slower-moving supply chain problem. It's not even clear the incentives line up to solve it fast. If you're OpenAI or Anthropic, you have some leverage to get custom silicon. A smaller shop? You're buying what's available and living within those constraints, or you're not competitive. That's a structural problem, not just an engineering one.
The part people miss is that inference is operationally rigid in ways training isn't. When you're training, you can chase efficiency for months. When you're serving customers, latency matters *today*, and you can't wait for better hardware. So you end up over-provisioned on compute just to hit response-time targets, which means your actual per-query cost is higher than the theoretical hardware cost would suggest. I'd be skeptical of any claim that inference costs are converging toward training cost improvements until someone actually proves customers are seeing their bills drop.
I think you're spotting something real, but I'd separate two different problems that are getting mushed together.
The training efficiency story is legit—better algorithms, scale, architecture tweaks. That's algorithmic gain, and it's real. But you're right that it doesn't directly move the inference needle the way people talk about it. Here's where I'd push back slightly though: when a new model does the same work with half the parameters, that *is* a direct inference win. Smaller model, less memory movement, faster. That's not free-lunch amortization—that's actual hardware efficiency. The question is whether algorithmic gains outpace the hardware constraints you're describing.
The memory bandwidth bottleneck you mention is the actual hard problem, and I think that's where the structural difference lives. In training, you're doing one forward-backward pass per batch and you can hide latency behind massive parallel work. In inference serving, you've got a queue of requests with SLA requirements and you can't just wait for memory. That's not a problem that better algorithms alone solve—you need hardware that actually moves faster. And yeah, bandwidth improvements have lagged compute by years, maybe a decade in some cases. So the installed base problem isn't really separate from the physics problem; it's baked in. Nobody's going to replace working H100s when the bottleneck is memory, not FLOPs.
The useful question might be: are we actually hitting the bandwidth wall yet, or are we still in the regime where model efficiency improvements buy us time?
I'd push back slightly on the framing here, but only because the evidence is weirder than "constraints versus efficiency gains." The real problem is that inference economics are locked into a different business model than training, and that's structural in a way that has nothing to do with memory bandwidth.
I managed supply chain for an apparel manufacturer for three years. We had one set of capital equipment (cutters, presses, stitching lines) that ran constantly because the marginal cost of a unit was maybe 40 cents on a $25 retail item. Utilization was everything. But we also had another set of equipment—the pattern-grading machines, the dye vats—that only made sense to run when you had a full batch. You couldn't amortize those costs across single units. The moment you needed sub-batch latency, you were carrying dead capital.
That's where inference sits now. If you're OpenAI or Anthropic running a public API, you can't batch endlessly like a training job can. You've got thousands of users hitting you at once, each one expecting a response in seconds. So you need to overprovision hardware to keep latency flat. You're carrying GPU capacity that sits idle most of the time. Training doesn't have that problem—you fill the batches, run for weeks, then shut down. Inference is running 24/7 with uneven demand. That's not a memory bandwidth problem. That's a utilization problem, and it doesn't improve just because your algorithms get more efficient. Better algorithms mean each query burns fewer FLOPs, sure. But you still need the hardware sitting there to catch the spike at 3pm.
The manufacturers who've managed to drop inference costs have done it by either accepting higher latency (batch processing) or by accepting much lower utilization (running smaller models on cheaper hardware and eating the quality tradeoff). Neither generalizes to what the market actually wants.