New Benchmark AA-AgentPerf: GB300 Hits 20x H200 Concurrent Agents per MW
新智元 · wechat · 2026-07-04
Independent evaluation agency ArtificialAnalysis released AA-AgentPerf, the industry's first inference benchmark designed specifically for AI agents, with "concurrent agents per megawatt" as its primary metric. Instead of feeding fixed-length synthetic requests, it replays real coding agent trajectories covering 12+ languages, up to 200 turns, and contexts reaching 130,000 tokens. It enables production optimizations like KV cache reuse and speculative decoding, locking in SLO service standards before testing maximum concurrency.
Nvidia's published AA-AgentPerf results show that under service standards of 20 and 60 tokens per second, the GB300NVL72's concurrent agents per megawatt are about 20 times those of the H200—for the same one megawatt of electricity, GB300NVL72 can handle roughly 61,400 concurrent agents, while H200 handles only about 2,600.
The article notes that agent workloads are relay-style long-chain calls; single-request benchmarks can no longer measure these chained invocations, tool waiting, and context bloat. For those spending real money on GPUs to build data centers, the real concern is how many working agents can be sustained per kilowatt-hour per GPU.
More from Infra
- Intel 10-Q points to 18A/14A progress and “potential significant external customers” — BenBajarin · 2026-07-27
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27