SemiAnalysis: 4-hi HBM wins on $/bandwidth, cutting inference cost amid DRAM shortage

dylan522p · x · 2026-09-14

SemiAnalysis argues the HBM stacking race is reversing as DRAM shortages bite: next-gen accelerators will standardize on 8-hi stacks instead of today's 12-hi — Nvidia's Rubin Ultra drops from 288GB to 192GB per GPU.

The deeper claim: for bandwidth-bound inference workloads, 4-hi HBM delivers the best $/bandwidth and thus the lowest cost per token. Beyond a capacity threshold, extra stacking offers diminishing returns while carrying the same BOM penalty. Hardware teams at major labs plan to adopt this approach from HBM4 in their ASIC programs, echoing the focus on maximizing tokens/Watt under datacenter power constraints.

Original post →

More from Infra

Infra channel →