Inference Bottleneck Evolution: HBM Capacity Now Limits Batch Size, Not KV Reads
YouJiacheng · x · 2026-09-11
YouJiacheng outlines three stages of LLM inference bottlenecks: early workloads were weight-loading bound; long-context workloads made KV cache reads the bottleneck; and now, with DSA and even longer contexts, batch size is constrained by HBM capacity instead. He argues this is driven by the nature of the model and workload, not something speculative decoding can change.
More from Infra
- OpenAI Agents API hits public beta; Cloudflare ships sandbox integration for cloud Codex agents — ritakozlov · 2026-09-11
- DeepSeek-V4.1-Flash lands on Fireworks: 552B MoE for coding and agents at 1/40th claimed cost — lqiao · 2026-09-11
- OpenAI Agents API Meets Cloudflare: Deploy a Logged Agent in 4 Minutes — craigsdennis · 2026-09-11
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11