vLLM thread (2/5): the two benchmarks behind the Vera Rubin numbers

vllm_project · x · 2026-10-10

Part 2 of vLLM's Vera Rubin thread details the benchmarks: 7.8x+ GB200 throughput on SemiAnalysis AgentX serving MiniMax M3 on Rubin NVL72 at matched interactivity, and up to 3.7x GB300 NVL72 throughput in MLPerf Inference v6.1 with Dynamo on Qwen3-VL-235B-A22B.

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →