vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3

vllm_project · x · 2026-10-10

The vLLM project announced early support for NVIDIA Vera Rubin, ported since the chip's reveal by vLLM community members alongside @inferact, NVIDIA, and Red Hat. On SemiAnalysis's AgentX benchmark, vLLM serving MiniMax M3 on a Vera Rubin NVL72 delivers 7.8x+ the throughput of GB200 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x GB300 NVL72 throughput on Qwen3-VL-235B-A22B.

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →