vLLM ships Vera Rubin NVL72 support with 7.8x per-GPU throughput over GB200

vllm_project · x · 2026-10-10

The vLLM team, with NVIDIA, Red Hat and Inferact, announced day-0 support for NVIDIA Vera Rubin NVL72, available today via nightly containers (vllm/vllm-openai:cu134-nightly) running DeepSeek, Kimi, GLM and MiniMax models.

Key points:

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →