FULL STORY
NVIDIA Showcases Vera Rubin Real-World Performance
NVIDIA first disclosed silicon-based real-world results for Vera Rubin NVL72, then cited SemiAnalysis' AgentX benchmark, both touting up to 30x gains over GB300 and rivals.
2026-08-24 ~ 2026-08-29 · 2 episodes · 7 posts
Episode 1 · NVIDIA Shows Vera Rubin Silicon: Up to 30x Agent Throughput over GB300 (2026-08-24, 5 posts)
NVIDIA published the first silicon-based performance results for Vera Rubin NVL72, claiming order-of-magnitude gains over GB300 NVL72 on real AI Agent workloads — the architecture's first official real-silicon validation aimed at the Agent era.
Confirmed
- The test was published on NVIDIA's technical blog, using SemiAnalysis's AgentX workload and the DeepSeek V4 Pro model against real agent traces.
- At matched interactivity targets, Vera Rubin delivers up to 30x higher throughput per megawatt and 35x lower agentic coding token cost than GB300 NVL72.
- The benchmark replays production-grade coding agent sessions, including long context and KV-cache reuse scenarios.
- The news was first posted on 08-24 by @nvidia and @nordicinst, then spread via reposts from @Scobleizer and @brianryhuang.
Why it matters
- This is NVIDIA's first demonstration of Rubin's capabilities from real silicon rather than paper specs, targeting AI agents — an emerging mainstream workload — instead of traditional training/inference benchmarks.
- Throughput-per-megawatt and token cost speak directly to data-center efficiency and unit economics of intelligence, offering a concrete reference for cloud providers and agent teams planning next-generation compute.
- As @Crescitaly notes, the figures are NVIDIA's own claims from its blog with no independent third-party replication yet; readers should mind the test conditions such as "matched interactivity targets."
Unconfirmed
- No independent replication exists yet; the 30x and 35x figures rest solely on NVIDIA's single-source methodology.
- NVIDIA Vera Rubin NVL72 delivers 30x higher efficiency for AI agents vs GB300 — nvidia · 2026-08-24
- Nvidia Vera Rubin NVL72 delivers 30x more work per watt than GB300 on agentic workloads — nordicinst · 2026-08-24
- NVIDIA Vera Rubin benchmarks show 35x cheaper agentic coding tokens — brianryhuang · 2026-08-25
- NVIDIA Vera Rubin benchmarks show 30x throughput boost for agent workloads — Scobleizer · 2026-08-25
- NVIDIA claims up to 30× agentic throughput per MW on Vera Rubin — Crescitaly · 2026-08-26
Episode 2 · NVIDIA Claims Rubin NVL72 Delivers 30x Better Efficiency on AgentX Benchmark (2026-08-29, 2 posts)
NVIDIA cites SemiAnalysis' AgentX, the first open-source benchmark for multi-turn agent workflows with up to 1M-token context, showing its Vera Rubin NVL72 delivers 30x higher throughput and energy efficiency than competitors.
- NVIDIA Cites SemiAnalysis AgentX: Rubin NVL72 Shows 30x Better Throughput — nvidia · 2026-08-29
- NVL72 Achieves Up to 30x Better Throughput per MW than GB300 on AgentX Benchmark — nvidia · 2026-08-29