Cutting 12-Hi to 8-Hi lifts bandwidth-per-GB 50%: the Rubin Ultra HBM math
AccBalanced · x · 2026-09-06
SemiAnalysis continues the math: capacity ≈ stacks × dies per stack × GB per die, while bandwidth ≈ stacks × interface width × pin speed. Stack height buys capacity, not bandwidth.
- Cutting 12-Hi to 8-Hi drops a third of the DRAM dies — the most supply-constrained silicon on the BOM — while bandwidth holds or ticks up
- Bandwidth per GB of capacity rises 50%, meaningfully improving $/bandwidth right as HBM prices hike
- Inference is bandwidth bound: every decoded token re-reads weights and the KV cache, so Nvidia's trade favors $/bandwidth over $/capacity
Related event: SemiAnalysis Explains Why Nvidia Cut Rubin Ultra HBM to 8-Hi(2 posts)→
More from Infra
- A First-Principles Handbook on KV Cache: From MHA/GQA/MLA to PagedAttention — techNmak · 2026-09-06
- The best local model you can run on 2 GB10s, per this desk setup — jasonkneen · 2026-09-06
- One ComfyUI node fixed MiniMax H3 OOM on RTX 5090: full workflow for 15s 2K video — denizbuyukayak · 2026-09-06
- KV cache often spills out of HBM in the agentic era, tanking effective bandwidth — AccBalanced · 2026-09-06
- Hybrid bonded HBM hypothetical market: over 3 billion D2D applications per year — zephyr_z9 · 2026-09-06
- Ollama CEO: open models will carry 80-90% of enterprise tokens at just 10-20% of cost — victor_explore · 2026-09-06