Vera Rubin NVL72 debuts in MLPerf v6.1 with up to 3.7x throughput vs GB300 NVL72
nordicinst · x · 2026-09-16
NVIDIA's Vera Rubin NVL72 made its first MLPerf Inference v6.1 preview submission, delivering up to 3.7x better throughput than GB300 NVL72. A 288-GPU GB300 submission hit 99% scaling efficiency, and software optimizations brought up to 1.6x gains over v6.0. NVIDIA stresses platform fungibility across models and workloads as the key to inference economics.
More from Infra
- Anthropic engineer's plea: 'I just want to serve five petaflops' — charles_irl · 2026-09-17
- Legora's legal AI search went from 100ms to 20s P99 — full infra postmortem — AI Engineer · 2026-09-17
- El Pulpo 0.1.0: A 12.7MB Proxy and Load Balancer for Local LLM Inference — zaytzev · 2026-09-16
- CPO summit takeaway: no single optical interconnect solution will win as AI scales — BenBajarin · 2026-09-16
- New arXiv Paper Proposes Intelligence per Watt: Measuring Efficiency of Local AI — yogthos · 2026-09-16
- Engineer breaks down X's search ranking stack: two-tower retrieval plus L1/L2 rankers — _jaydeepkarale · 2026-09-16