SemiAnalysis Open Sources $3M AgentX Benchmark for Agentic Coding Workloads

AccBalanced · x · 2026-08-24

SemiAnalysis released AgentX 1.0, the first open-source, multi-turn agentic coding inference benchmark at 1 million context, built with over $3M.

It addresses the inaccuracy of traditional fixed-sequence benchmarks by modeling realistic agentic workflows: multi-turn, long context, high prefill reuse, and sub-agent bursts. The test matrix spans 2MW of compute across chips like GB300 NVL72, MI355, and B200.

Original post →

More from coding & agent

coding & agent channel →