Why HBM stacks are diluted bandwidth: 20 layers leaves you at 20% of single-layer speed, slower than DDR5
knowrohit07 · x · 2026-09-25
An engineer argues HBM stacking is the wrong path: a single DRAM layer has 20 TB/cm² of cell-level bandwidth (derivable from refresh calculations). Stack 20 layers (roughly double HBM4) and each layer gets only 4 TB/cm² — 20% of single-layer throughput, hugely diluted. At 20 high, average HBM4E speed is slower than DDR5. He questions why SK Hynix and Samsung prioritize height over speed instead of intelligently placing cheaper memory around the stack.
Related event: Engineer Questions HBM Stacking: Bandwidth Diluted to 20% at 20 Layers(4 posts)→
More from Infra
- KV Cache Explained: How a Simple Idea Turns Inference Into Systems Engineering — Abhishekcur · 2026-09-25
- $360 for 160GB VRAM: building a local LLM server on unlocked CMP 50HX mining cards — Boricua-vet · 2026-09-25
- Former Intel CEO calls HBM "lousy" at Hot Chips 2026 as High Bandwidth Flash looms — Glittering_Depth_722 · 2026-09-25
- MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model — TheKanter · 2026-09-25
- kvcached brings virtual memory to LLM KV cache, deployed on 10K+ GPUs — techNmak · 2026-09-25
- IEEE plenary talk: micro-optimizations across the full stack, from silicon to models — fooobar · 2026-09-25