KV cache often spills out of HBM in the agentic era, tanking effective bandwidth
AccBalanced · x · 2026-09-06
A rebuttal to the view that inference performance tracks HBM pin bandwidth alone:
- The bandwidth argument only holds when the KV cache always fits in HBM. In the agentic era of long contexts and multi-turn tasks, it often doesn't.
- Once capacity runs out, tokens offload to slower DRAM/NAND tiers and latency spikes.
- In a tiered memory system, capacity and bandwidth are inseparable: HBM is the fast lane, and capacity is what keeps it from clogging. Cutting HBM capacity sacrifices the very bandwidth you thought you saved.
More from Infra
- Hybrid bonded HBM hypothetical market: over 3 billion D2D applications per year — zephyr_z9 · 2026-09-06
- Ollama CEO: open models will carry 80-90% of enterprise tokens at just 10-20% of cost — victor_explore · 2026-09-06
- Nvidia de-specced Rubin Ultra HBM from 12-Hi to 8-Hi: $/bandwidth is the bottleneck — AccBalanced · 2026-09-06
- A GPU running 5% slow is fine for inference but catastrophic for training: why health checks invert — AccBalanced · 2026-09-06
- Hot Chips 2026: Irrational Analysis publishes investment-driven recap — jwt0625 · 2026-09-06
- New method predicts transformer training divergence before the run starts — burkov · 2026-09-06