NVIDIA GB300 System Delivers 10x Inference Efficiency of Hopper
nvidia · x · 2026-07-15
NVIDIA officially emphasized that from benchmarks to actual production deployment, the GB300 NVL72 system delivers industry-leading performance per watt, designed to maximize revenue and achieve the lowest token costs.
In their responses, they revealed specific comparison data: when running the Kimi K2.6 model, the performance per watt of the GB300 NVL72 system can reach up to 10x that of the previous-generation Hopper architecture.
Related event: NVIDIA GB300 Delivers 10x Energy Efficiency Over Hopper(2 posts)→
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11