SemiAnalysis: 4-hi HBM wins on $/bandwidth, cutting inference cost amid DRAM shortage
dylan522p · x · 2026-09-14
SemiAnalysis argues the HBM stacking race is reversing as DRAM shortages bite: next-gen accelerators will standardize on 8-hi stacks instead of today's 12-hi — Nvidia's Rubin Ultra drops from 288GB to 192GB per GPU.
The deeper claim: for bandwidth-bound inference workloads, 4-hi HBM delivers the best $/bandwidth and thus the lowest cost per token. Beyond a capacity threshold, extra stacking offers diminishing returns while carrying the same BOM penalty. Hardware teams at major labs plan to adopt this approach from HBM4 in their ASIC programs, echoing the focus on maximizing tokens/Watt under datacenter power constraints.
More from Infra
- What if all AI API credits were condensed into a single tokenized currency? — sull · 2026-09-14
- Running Qwen 27B on 2x P40s: parallel agents wreck KV cache, seeking a serial agent harness — Jumpy-Operation-4615 · 2026-09-14
- TSMC paces the frontier by accident whenever it underestimates chip demand — dan_s_becker · 2026-09-14
- Dev warns: shady cheap-token inference providers send fake tool calls and resell your traces — Nils_Reimers · 2026-09-14
- Used RTX 5090 listed at £3,900 (~$5,200), more than double its MSRP — julianharris · 2026-09-14
- Why Amazon and Microsoft Are Taking Communities' Side Against Utilities — pstAsiatech · 2026-09-14