Inference Bottleneck Evolution: HBM Capacity Now Limits Batch Size, Not KV Reads

YouJiacheng · x · 2026-09-11

YouJiacheng outlines three stages of LLM inference bottlenecks: early workloads were weight-loading bound; long-context workloads made KV cache reads the bottleneck; and now, with DSA and even longer contexts, batch size is constrained by HBM capacity instead. He argues this is driven by the nature of the model and workload, not something speculative decoding can change.

Related event: How LLM Inference Bottlenecks Evolved: From Weights to KV Reads to HBM Capacity(2 posts)→

Original post →

More from Infra

Infra channel →