vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput
vLLM officially announced support for the NVIDIA Vera Rubin NVL72. Since Rubin's launch, the community along with Inferact, NVIDIA, and Red Hat have been pushing the port forward, and a daily-built container image (vllm/vllm-openai:cu134-nightly) is already available and runnable today.
Confirmed
- On the SemiAnalysis AgentX benchmark, vLLM serving MiniMax M3 on Vera Rubin NVL72 delivered over 7.8x the throughput of GB200 at equivalent interactivity; MLPerf Inference results were also shared (in part two of vLLM's official thread).
- Hardware comparison (from vLLM's official thread): vs. GB200 NVL72, 3.5x FP4 FLOPS, roughly 2.4x HBM bandwidth, 1.7x NVLink bandwidth, and 2-4x faster softmax exponentials; the stronger NVLink cuts MoE (expert parallelism) communication overhead.
- SemiAnalysis analysis says the Rubin platform delivers 3.2x the profit per gigawatt of GB300 NVL72 on production inference engine vLLM, and up to 10x performance per dollar.
Why it matters
- Rubin NVL72 is positioned as the next-gen platform for agentic inference; vLLM completing support this early with reproducible benchmark data shows the open-source inference stack is keeping pace with new hardware at a markedly faster cadence.
- SemiAnalysis's per-gigawatt profit and per-dollar performance metrics offer quantified references for datacenter-level procurement and deployment decisions.
2026-10-10 ~ 2026-10-10 · 7 related posts
Primary sources
- SemiAnalysis: NVIDIA Rubin with vLLM Delivers 3.2x Profit per Gigawatt, up to 10x Perf per Dollar vs GB300 NVL72 — woosuk_k · 2026-10-10
- [source] vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3 — vllm_project · 2026-10-10
- vLLM thread (2/5): the two benchmarks behind the Vera Rubin numbers — vllm_project · 2026-10-10
- [source] vLLM thread (3/5): Rubin vs GB200 — 3.5x FP4 FLOPS, 2.4x HBM bandwidth — vllm_project · 2026-10-10
- [source] vLLM ships Vera Rubin NVL72 support with 7.8x per-GPU throughput over GB200 — vllm_project · 2026-10-10
2 near-duplicate retellings: Xianbao_QIAN · woosuk_k