vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput

vLLM officially announced support for the NVIDIA Vera Rubin NVL72. Since Rubin's launch, the community along with Inferact, NVIDIA, and Red Hat have been pushing the port forward, and a daily-built container image (vllm/vllm-openai:cu134-nightly) is already available and runnable today.

Confirmed

Why it matters

2026-10-10 ~ 2026-10-10 · 7 related posts

Primary sources

2 near-duplicate retellings: Xianbao_QIAN · woosuk_k