Vera Rubin NVL72 Delivers 5x Inference Performance Per Watt Over GB200
downingARK · x · 2026-07-30
SemiAnalysis released an in-depth inference TCO and architecture analysis comparing NVIDIA's next-gen Vera Rubin NVL72 with the current GB200 NVL72.
- Performance Leap: Early engineering samples show Vera Rubin delivering 5.4x performance per MW and 5x performance per dollar over GB200 NVL72 when running DeepSeek R1. As Rubin is still in early bring-up, this gap is expected to widen.
- Smoother Software Transition: NVIDIA has released the first Rubin (SM107) software stack with CUDA 13.4 and upstreamed PRs to PyTorch, vLLM, and OpenAI Triton. Unlike the Hopper-to-Blackwell transition, Rubin can reuse Blackwell's WGMMA kernels, significantly accelerating time-to-market.
- Architecture Details: The analysis details Rubin's new 3-bit programmable LUT tensor core. Additionally, GitHub info reveals the next-gen Feynman architecture is SM140, and the Rubin-to-Feynman kernel transition will be substantially more complex.
More from Infra
- Cloud Revenue Growth Showdown: Google Cloud Hits 63% in Q1 2026, Beating Azure and AWS — Beth_Kindig · 2026-07-30
- 1-bit Quantization Magic: Kimi K3 Successfully Runs Locally on Mac Studio — danielhanchen · 2026-07-30
- LMSYS Debuts Miles: Blackwell-Native 8-bit and 4-bit RL Recipes — BanghuaZ · 2026-07-30
- llama.cpp Update: Default MTP Tensor Loading Increases VRAM Usage — Shoddy_Bed3240 · 2026-07-30
- Open Community Boosts Local LLM Inference on Mac by 80.6% Without Speculative Decoding — gajesh · 2026-07-30
- Kimi K3 runs on vLLM + AMD from Day 0, supporting 2.8T params on Instinct — vllm_project · 2026-07-30