SemiAnalysis: 4-hi HBM wins on $/bandwidth as Rubin Ultra drops to 192GB 8-hi stacks
maksym_andr · x · 2026-09-15
SemiAnalysis argues the ever-taller HBM stacking trend is breaking.
- Next-gen accelerators standardize on 8-hi instead of 12-hi: Nvidia's Rubin Ultra drops to 192GB from 288GB per GPU (SemiAnalysis first reported), and the supply chain is prepping 8-hi even though 16-hi was the expectation under a year ago;
- For inference workloads where bandwidth matters most, 4-hi HBM offers the best $/bandwidth and lowest cost per token; beyond a capacity threshold, extra bits carry BOM penalty with diminishing returns;
- Major labs' ASIC teams want this from HBM4 onward, alongside maximizing tokens/Watt under DC power constraints.
The shift also hints at relief for today's extreme DRAM shortage.
Related event: SemiAnalysis: HBM Stacking Trend Reverses Toward 8-hi(2 posts)→
More from Infra
- NVIDIA: full-stack NIM tuning delivers 2.5x more concurrent users on Nemotron 3 Ultra — NVIDIAAI · 2026-09-15
- How Much Does Local LLM Inference Really Cost? A Dev Added an Electricity Calculator — giveen · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- Single Pure-C99 Inference Engine Runs Both BitNet Ternary and GGUF, No Python or CUDA — shifu_legend · 2026-09-15
- Dev weighs ChatGPT subscription via OAuth vs API pricing for a production RAG app — builtforoutput · 2026-09-15
- Cognichip launches ACI Enterprise: one engineer finishes chip front-end design in 10 days — kimmonismus · 2026-09-15