NVIDIA claims up to 30× agentic throughput per MW on Vera Rubin

Crescitaly · reddit · 2026-08-26

NVIDIA's technical blog showcases AgentX results on the Vera Rubin architecture, claiming up to 30× higher throughput per megawatt than GB300 at a matched interactivity target. The test replays production-style coding agent sessions with long context and KV-cache reuse. The author argues that while more representative than fixed context serving, focusing solely on token output doesn't equal useful work, suggesting production benchmarks should include task outcomes, tool amplification, cache hit rates, and rollback rates.

Related event: NVIDIA Shows Vera Rubin Delivering Up to 30x Agent Throughput over GB300(5 posts)→

Original post →

More from coding & agent

coding & agent channel →