SemiAnalysis: Nvidia's custom NVHBM frees ~25% more compute die area on Feynman
zephyr_z9 · x · 2026-10-03
SemiAnalysis breaks down Nvidia's custom HBM approach, NVHBM. On Rubin, HBM controllers and PHYs take roughly 16% of the GPU die; on Feynman with NVHBM, that's estimated to drop to 4%.
The trick: the memory controller moves off the XPU onto the HBM base die, and the wide standard PHY is replaced by Nvidia's compact NV-HBI die-to-die link. Samsung's custom HBM4 interface is 60% smaller than its standard PHY.
Nvidia claims up to 25% more die area for compute, up to 30% more bandwidth, and 15% lower HBM power versus standard HBM4E.
Related event: SemiAnalysis: Custom NVHBM Could Free Up ~25% of Die Area on Nvidia Feynman(2 posts)→
More from Infra
- Qwen 177B at 11-15 tok/s on a Single RTX 5070 12GB: Expert Streaming Deep Dive — ayobluestarr · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- llama.cpp PR Halves Indexer Score Memory for Qwen Flash, Cutting VRAM Use — jacek2023 · 2026-10-03
- gufo-Qwen3.6-35B hits 3095 tok/s prefill, 190 tok/s decode on Strix Halo — nubela · 2026-10-03
- Garage server farms return: GPU and power shortages reverse the AWS era in Palo Alto — bookwormengr · 2026-10-03
- AI maxi calls GPU price hike a bubble peak, plans to buy cheap cards after burst — AIFlow_ML · 2026-10-03