FULL STORY

NVIDIA Showcases Vera Rubin Real-World Performance

NVIDIA first disclosed silicon-based real-world results for Vera Rubin NVL72, then cited SemiAnalysis' AgentX benchmark, both touting up to 30x gains over GB300 and rivals.

2026-08-24 ~ 2026-08-29 · 2 episodes · 7 posts

Episode 1 · NVIDIA Shows Vera Rubin Silicon: Up to 30x Agent Throughput over GB300 (2026-08-24, 5 posts)

NVIDIA published the first silicon-based performance results for Vera Rubin NVL72, claiming order-of-magnitude gains over GB300 NVL72 on real AI Agent workloads — the architecture's first official real-silicon validation aimed at the Agent era.

Confirmed

  • The test was published on NVIDIA's technical blog, using SemiAnalysis's AgentX workload and the DeepSeek V4 Pro model against real agent traces.
  • At matched interactivity targets, Vera Rubin delivers up to 30x higher throughput per megawatt and 35x lower agentic coding token cost than GB300 NVL72.
  • The benchmark replays production-grade coding agent sessions, including long context and KV-cache reuse scenarios.
  • The news was first posted on 08-24 by @nvidia and @nordicinst, then spread via reposts from @Scobleizer and @brianryhuang.

Why it matters

  • This is NVIDIA's first demonstration of Rubin's capabilities from real silicon rather than paper specs, targeting AI agents — an emerging mainstream workload — instead of traditional training/inference benchmarks.
  • Throughput-per-megawatt and token cost speak directly to data-center efficiency and unit economics of intelligence, offering a concrete reference for cloud providers and agent teams planning next-generation compute.
  • As @Crescitaly notes, the figures are NVIDIA's own claims from its blog with no independent third-party replication yet; readers should mind the test conditions such as "matched interactivity targets."

Unconfirmed

  • No independent replication exists yet; the 30x and 35x figures rest solely on NVIDIA's single-source methodology.

Episode 2 · NVIDIA Claims Rubin NVL72 Delivers 30x Better Efficiency on AgentX Benchmark (2026-08-29, 2 posts)

NVIDIA cites SemiAnalysis' AgentX, the first open-source benchmark for multi-turn agent workflows with up to 1M-token context, showing its Vera Rubin NVL72 delivers 30x higher throughput and energy efficiency than competitors.