vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3
vllm_project · x · 2026-10-10
The vLLM project announced early support for NVIDIA Vera Rubin, ported since the chip's reveal by vLLM community members alongside @inferact, NVIDIA, and Red Hat. On SemiAnalysis's AgentX benchmark, vLLM serving MiniMax M3 on a Vera Rubin NVL72 delivers 7.8x+ the throughput of GB200 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x GB300 NVL72 throughput on Qwen3-VL-235B-A22B.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- Starship Could Cut Cost to Orbit 100x to ~$185K Per Ton, Unlocking New Businesses — claud_fuen · 2026-10-10
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- Nvidia CEO's son-in-law becomes VP as Apple cuts iPhone 18 Pro orders 15%+ — 创业邦 · 2026-10-10