vLLM thread (3/5): Rubin vs GB200 — 3.5x FP4 FLOPS, 2.4x HBM bandwidth
vllm_project · x · 2026-10-10
Part 3 of vLLM's thread compares Rubin to GB200: 3.5x FP4 FLOPS, 2.4x HBM bandwidth, 1.7x NVLink bandwidth, and 2-4x faster softmax exponentials; faster NVLink cuts AllReduce and all-to-all costs in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, while FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.
Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→
More from Infra
- US grid adds 86GW this year while AI labs need hundreds of GW of power — FinanceYF5 · 2026-10-10
- SpaceX building industrial base to make up to 1,000 Starships a year, with Florida GigaBay 11x bigger than current Megabay — XFreeze · 2026-10-10
- One CUDA Guide chapter beats 96% of people on GPU execution and memory management — blelbach · 2026-10-10
- Fireworks AI confirms security incident involving unauthorized use of internal credentials — lqiao · 2026-10-10
- ChapterPal Brings Offline Gemini Nano AI Tutor to Android, iOS Version Coming Soon — burkov · 2026-10-10
- Chat with AI could feel outdated in 1-2 years as agent swarms push traffic 1,000x — dumpshoot · 2026-10-10