Memory wall chart: compute up ~3x per two years, HBM bandwidth under 2x
Summit-Star001 · reddit · 2026-09-14
From Raghu Sreeramanenei's memory tutorial at Hot Chips 2026: normalized compute (TPU v3 through R200) grows 3x every two years while HBM bandwidth (HBM2e through HBM4) grows under 2x — and the log scale hides the widening gap. Three fixes are in flight: memory beside compute, memory closer on shorter links, and multiply units inside memory itself. Samsung's approach already ships in LPDDR5X, measuring 3.01x tokens/sec on Llama 3.1 8B.
More from Infra
- Perplexity launches Hybrid Compute to split AI tasks between cloud and local Mac — Aiden_Tech_Ai · 2026-09-14
- 4 sink tokens + 64-token window matches distilled linear attention, no training needed — burny_tech · 2026-09-14
- RDNA4 local inference hits ~100 tok/s running Qwen3.8 Flash on dual R9700 — Public_Umpire_1099 · 2026-09-14
- Anthropic reportedly signed a $13.7B, 6-year compute deal with RUM Group's Georgia site — rohanpaul_ai · 2026-09-14
- AI-written OpenSCAD + dual 12-inch fans fix NVIDIA Thor Dev Kit thermal throttling — catplusplusok · 2026-09-14
- Data centers are crowding out US private construction, WSJ data shows — GregCook2011 · 2026-09-14