NVIDIA Vera Rubin NVL72 delivers 30x higher efficiency for AI agents vs GB300

nvidia · x · 2026-08-24

NVIDIA released the first on-silicon performance data for Vera Rubin, measured on real-world agentic workloads using the DeepSeek V4 Pro model and SemiAnalysis AgentX benchmark. Results show Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt and 35x lower token costs compared to GB300 NVL72. The tests highlight that agentic sessions involve long-context growth across hundreds of steps, differing significantly from chat or summarization tasks.

Related event: NVIDIA's Vera Rubin Delivers 30x Gains Over GB300 in Agent Workloads(2 posts)→

Original post →

More from Infra

Infra channel →