NVIDIA Vera Rubin benchmarks show 35x cheaper agentic coding tokens
brianryhuang · x · 2026-08-25
NVIDIA released the first on-silicon benchmarks for the Vera Rubin NVL72, measured against the SemiAnalysis AgentX workload using the DeepSeek V4 Pro model. The results show up to 30x higher throughput per megawatt and 35x lower token costs compared to GB300 NVL72.
Key drivers include:
- Full-stack Co-design: Enhanced TensorCore FLOPs, HBM4 bandwidth, 6th Gen NVLink/Spectrum-6 SPX networking, and Vera CPUs for faster tool calls.
- Disaggregated Serving: Separating prefill and decode with rate matching, optimized for the Rubin+LPX architecture.
- Expert Parallelism: Spreading experts across many GPUs to scale capacity.
Related event: NVIDIA's Rubin Benchmarks Show 30x Gains Over GB300 for AI Agents(3 posts)→
More from coding & agent
- Pi Contributor Debunks 'Official Critique' as Misread Unreviewed PR — dotey · 2026-08-25
- Agent Bottlenecks Shift to Overhead: Tool Calls and I/O Lag Inference Speed — MatthewBerman · 2026-08-25
- Dev Choice Shift: From DeepSeek Harness to Pi for Stability — dotey · 2026-08-25
- Grok Build Adds 'Browser Use' Plugin for Local Chrome and Cloud Browsing — elonmusk · 2026-08-25
- Grok Review: Independent VMs and Browser Access Enable Seamless Multi-Bot Workflows — brandon_galang · 2026-08-25
- Survey on Terminal Agents: Definitions and Evaluation Frameworks — omarsar0 · 2026-08-25