NVIDIA Launches Vera Rubin with 10x Energy Efficiency

NVIDIA has officially launched the Vera Rubin platform, designed for the Agentic era to significantly reduce inference costs for large-scale AI models by maximizing compute efficiency. The Vera Rubin NVL72 is currently being rolled out with global partners including CoreWeave, Google Cloud, and Microsoft.

Confirmed

NVIDIA claims that Vera Rubin delivers 10x better performance per watt. According to initial real-world performance data obtained by CoreWeave, the Vera Rubin NVL72 achieves 10x more token generation per megawatt compared to existing Blackwell architectures (like the GB200) when running the DeepSeek-R1 model. Furthermore, NVIDIA executive Kyle Kranen disclosed the design goals for the next-generation Rubin architecture, aiming to maximize "agentic tokens/watt." The core hardware specifications include a compute target of 50 Petaflops and HBM4 memory with a bandwidth of 22TB/s.

Why it matters

As AI models scale and agentic applications become widespread, the compute and energy costs during inference have become a primary bottleneck. By delivering a 10x improvement in performance per watt, Vera Rubin is poised to drastically reduce the operational costs of large-scale AI and accelerate the adoption of highly energy-efficient infrastructure.

2026-07-21 ~ 2026-07-23 · 6 related posts

Full story(4 episodes)→

Primary sources

1 near-duplicate retellings: nvidia