vLLM thread (3/5): Rubin vs GB200 — 3.5x FP4 FLOPS, 2.4x HBM bandwidth

vllm_project · x · 2026-10-10

Part 3 of vLLM's thread compares Rubin to GB200: 3.5x FP4 FLOPS, 2.4x HBM bandwidth, 1.7x NVLink bandwidth, and 2-4x faster softmax exponentials; faster NVLink cuts AllReduce and all-to-all costs in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, while FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

Related event: vLLM Adds NVIDIA Vera Rubin NVL72 Support, Delivering 7.8x GB200 Throughput(7 posts)→

Original post →

More from Infra

Infra channel →