vLLM runs on NVIDIA Vera Rubin NVL72: 7.8x per-GPU throughput over GB200 NVL72

Xianbao_QIAN · x · 2026-10-10

The vLLM team, with NVIDIA, Red Hat, and Inferact, has brought up support for NVIDIA Vera Rubin NVL72. The platform offers 5x NVFP4 FLOPS, 2.4x HBM bandwidth, and 1.7x NVLink bandwidth versus GB200 NVL72. Day-0 support works via Blackwell-compatible kernels (DeepSeek, Kimi, GLM, MiniMax), plus Rubin-tuned FlashInfer 0.7.0 kernels and locality-aware MoE weight splitting using CUDA 13.4 locality domains. Early results: 7.8x per-GPU throughput over GB200 NVL72 on AgentX at matched interactivity, and up to 3.7x VLM throughput versus GB300 NVL72 in MLPerf.

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →