vLLM runs on NVIDIA Vera Rubin NVL72: 7.8x per-GPU throughput over GB200 NVL72
Xianbao_QIAN · x · 2026-10-10
The vLLM team, with NVIDIA, Red Hat, and Inferact, has brought up support for NVIDIA Vera Rubin NVL72. The platform offers 5x NVFP4 FLOPS, 2.4x HBM bandwidth, and 1.7x NVLink bandwidth versus GB200 NVL72. Day-0 support works via Blackwell-compatible kernels (DeepSeek, Kimi, GLM, MiniMax), plus Rubin-tuned FlashInfer 0.7.0 kernels and locality-aware MoE weight splitting using CUDA 13.4 locality domains. Early results: 7.8x per-GPU throughput over GB200 NVL72 on AgentX at matched interactivity, and up to 3.7x VLM throughput versus GB300 NVL72 in MLPerf.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- Building a pit crew for Grok Bot: frontier model plans, free models grind — alexcovo_eth · 2026-10-10
- AI boom turns into a debt boom: Oracle 5y CDS near record 261bps, implying 20.4% default odds — cyb3rops · 2026-10-10
- Fireworks AI Discloses Security Incident Involving Unauthorized Use of Internal Credentials — lqiao · 2026-10-10
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- Running Qwen3.6 35B-A3B with 131K context and vision on a 6GB RTX 2060 — full config — Szadbaverem69 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10