vLLM thread (2/5): the two benchmarks behind the Vera Rubin numbers
vllm_project · x · 2026-10-10
Part 2 of vLLM's Vera Rubin thread details the benchmarks: 7.8x+ GB200 throughput on SemiAnalysis AgentX serving MiniMax M3 on Rubin NVL72 at matched interactivity, and up to 3.7x GB300 NVL72 throughput in MLPerf Inference v6.1 with Dynamo on Qwen3-VL-235B-A22B.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10
- One CUDA Guide chapter beats 96% of people on GPU execution and memory management — blelbach · 2026-10-10
- Fireworks AI confirms security incident involving unauthorized use of internal credentials — lqiao · 2026-10-10
- ChapterPal Brings Offline Gemini Nano AI Tutor to Android, iOS Version Coming Soon — burkov · 2026-10-10
- Chat with AI could feel outdated in 1-2 years as agent swarms push traffic 1,000x — dumpshoot · 2026-10-10