New Benchmark AA-AgentPerf: GB300 Hits 20x H200 Concurrent Agents per MW
新智元 · wechat · 2026-07-04
Independent evaluation agency ArtificialAnalysis released AA-AgentPerf, the industry's first inference benchmark designed specifically for AI agents, with "concurrent agents per megawatt" as its primary metric. Instead of feeding fixed-length synthetic requests, it replays real coding agent trajectories covering 12+ languages, up to 200 turns, and contexts reaching 130,000 tokens. It enables production optimizations like KV cache reuse and speculative decoding, locking in SLO service standards before testing maximum concurrency.
Nvidia's published AA-AgentPerf results show that under service standards of 20 and 60 tokens per second, the GB300NVL72's concurrent agents per megawatt are about 20 times those of the H200—for the same one megawatt of electricity, GB300NVL72 can handle roughly 61,400 concurrent agents, while H200 handles only about 2,600.
The article notes that agent workloads are relay-style long-chain calls; single-request benchmarks can no longer measure these chained invocations, tool waiting, and context bloat. For those spending real money on GPUs to build data centers, the real concern is how many working agents can be sustained per kilowatt-hour per GPU.
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11