NVIDIA Vera Rubin NVL72 Delivers 30x Efficiency Gain for Agentic AI
NVIDIA Blog · rss · 2026-08-24
NVIDIA released data showing its Vera Rubin NVL72 system delivers up to 30x higher throughput per megawatt and 35x lower cost per million tokens than GB300 NVL72 on agentic workloads. Based on SemiAnalysis AgentX real-world trajectories, Vera Rubin utilizes extreme codesign techniques like disaggregated serving, rate matching, and distributed KV-caching to boost efficiency for long-context and tool-intensive tasks.
More from Infra
- Opinion: 'Prefill' Sounds Advanced But Is Simple Once You Understand LLMs — brandon_xyzw · 2026-08-27
- Nvidia's financials are insane; SpaceX may beat them in the future — mitchdeg · 2026-08-27
- Blue-Green Deployment Strategy for Zero-Downtime — _jaydeepkarale · 2026-08-27
- Chinese open models on Huawei chips said to crush US closed models on cost — chris_j_paxton · 2026-08-27
- Video generation speeds: 23.7s vs 11m shows Jevons Paradox in action — gorkem · 2026-08-27
- The Guardian podcast: Everyone hates datacentres, but do we really need them? — nordicinst · 2026-08-27