vLLM on NVIDIA Vera Rubin NVL72 hits 7.8x GB200 throughput on MiniMax M3
woosuk_k · x · 2026-10-10
vLLM announced early support for NVIDIA Vera Rubin: at matched interactivity, inference throughput on a Vera Rubin NVL72 running MiniMax M3 exceeds GB200 by more than 7.8x, measured on the AgentX benchmark. The Inferact team has co-led the effort since Rubin was announced, integrating a Rubin-optimized MSA prefill kernel and applying locality-aware tuning for Rubin GPUs. Results are early but signal a major generational jump in inference economics.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10
- One CUDA Guide chapter beats 96% of people on GPU execution and memory management — blelbach · 2026-10-10
- Fireworks AI confirms security incident involving unauthorized use of internal credentials — lqiao · 2026-10-10
- ChapterPal Brings Offline Gemini Nano AI Tutor to Android, iOS Version Coming Soon — burkov · 2026-10-10
- Chat with AI could feel outdated in 1-2 years as agent swarms push traffic 1,000x — dumpshoot · 2026-10-10