NVIDIA Vera Rubin NVL72 delivers 30x higher efficiency for AI agents vs GB300
nvidia · x · 2026-08-24
NVIDIA released the first on-silicon performance data for Vera Rubin, measured on real-world agentic workloads using the DeepSeek V4 Pro model and SemiAnalysis AgentX benchmark. Results show Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt and 35x lower token costs compared to GB300 NVL72. The tests highlight that agentic sessions involve long-context growth across hundreds of steps, differing significantly from chat or summarization tasks.
Related event: NVIDIA's Vera Rubin Delivers 30x Gains Over GB300 in Agent Workloads(2 posts)→
More from Infra
- Stripe to Acquire OpenRouter as Token Usage Hits 4.5+ Quadrillion Annual Rate — rohanpaul_ai · 2026-08-25
- OpenRouter usage surges 9,000x; agent workloads dominate token consumption — rohanpaul_ai · 2026-08-25
- Hot Chips conference main event starting in 12 minutes — firstadopter · 2026-08-25
- Teutonic-II-110B Announced: Exploring Loss-based Markets for Decentralized Training — markjeffrey · 2026-08-24
- Opinion: Nvidia + Open Ecosystem is the real Rebel Alliance — annbordetsky · 2026-08-24
- AI-Powered Metadata Correction and Harmonization on AWS — AWS ML Blog · 2026-08-24