Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x throughput vs GB300
NVIDIA Blog · rss · 2026-09-16
NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1: up to 3.7x throughput vs GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1, using NVFP4, disaggregated serving and expert parallelism. GB300 showed 99% scaling efficiency at 288 GPUs; software gains hit 1.6x over v6.0; 30x preview on SemiAnalysis AgentX.
More from Infra
- Crusoe runs 512 AMD MI355X GPUs at 5.75M tok/s in largest MLPerf inference entry — wkmyrhang · 2026-09-17
- Perovskite could lift solar efficiency ceiling from 30% to 45% — and give the US a shot against China — kyliebytes · 2026-09-17
- How mobile and specialization broke homogeneous compute into TPUs, NPUs, and more — blelbach · 2026-09-17
- After the x86 Monoculture: Software Will Suffer for Hardware's Fragmentation Again — blelbach · 2026-09-17
- Hardware veteran: low-precision gains nearly exhausted, true sparsity is AI's next 10x — blelbach · 2026-09-17
- Moore's Law in three eras: from free lunch (1970-2005) to software hell (2015-now) — blelbach · 2026-09-17