vLLM ships Vera Rubin NVL72 support with 7.8x per-GPU throughput over GB200
vllm_project · x · 2026-10-10
The vLLM team, with NVIDIA, Red Hat and Inferact, announced day-0 support for NVIDIA Vera Rubin NVL72, available today via nightly containers (vllm/vllm-openai:cu134-nightly) running DeepSeek, Kimi, GLM and MiniMax models.
Key points:
- Hardware: Vera Rubin NVL72 offers 5x the NVFP4 FLOPS, 2.4x HBM bandwidth and 1.7x bidirectional NVLink bandwidth of GB200 NVL72, plus 2-4x faster exponentials for softmax.
- Day-0 support: Rubin stays in the Blackwell architecture family, so vLLM's Blackwell kernels carry over directly.
- Tuned kernels: FlashInfer 0.7.0 brings Rubin-tuned attention, GEMM and MoE kernels, with the MiniMax Sparse Attention prefill kernel specifically tuned.
- Locality-aware MoE: Using CUDA 13.4 locality domains, MoE FC1/FC2 weights are sharded column-wise and launched per-domain via Green Contexts, so SMs read from the nearest HBM partition.
- Early performance: 7.8x per-GPU throughput versus GB200 NVL72 on AgentX at matched interactivity, and up to 3.7x higher VLM throughput versus GB300 NVL72 in MLPerf, with further optimizations expected.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10
- One CUDA Guide chapter beats 96% of people on GPU execution and memory management — blelbach · 2026-10-10
- Fireworks AI confirms security incident involving unauthorized use of internal credentials — lqiao · 2026-10-10
- ChapterPal Brings Offline Gemini Nano AI Tutor to Android, iOS Version Coming Soon — burkov · 2026-10-10
- Chat with AI could feel outdated in 1-2 years as agent swarms push traffic 1,000x — dumpshoot · 2026-10-10