Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x throughput vs GB300

NVIDIA Blog · rss · 2026-09-16

NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1: up to 3.7x throughput vs GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1, using NVFP4, disaggregated serving and expert parallelism. GB300 showed 99% scaling efficiency at 288 GPUs; software gains hit 1.6x over v6.0; 30x preview on SemiAnalysis AgentX.

Original post →

More from Infra

Infra channel →