vLLM on NVIDIA Vera Rubin NVL72 hits 7.8x GB200 throughput on MiniMax M3

woosuk_k · x · 2026-10-10

vLLM announced early support for NVIDIA Vera Rubin: at matched interactivity, inference throughput on a Vera Rubin NVL72 running MiniMax M3 exceeds GB200 by more than 7.8x, measured on the AgentX benchmark. The Inferact team has co-led the effort since Rubin was announced, integrating a Rubin-optimized MSA prefill kernel and applying locality-aware tuning for Rubin GPUs. Results are early but signal a major generational jump in inference economics.

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →