SemiAnalysis Open Sources $3M AgentX Benchmark for Agentic Coding Workloads
AccBalanced · x · 2026-08-24
SemiAnalysis released AgentX 1.0, the first open-source, multi-turn agentic coding inference benchmark at 1 million context, built with over $3M.
It addresses the inaccuracy of traditional fixed-sequence benchmarks by modeling realistic agentic workflows: multi-turn, long context, high prefill reuse, and sub-agent bursts. The test matrix spans 2MW of compute across chips like GB300 NVL72, MI355, and B200.
More from coding & agent
- Prompt Engineering: Enforce Single-Turn Questions to Prevent Info Overload — dejanseo · 2026-08-24
- More memory made my AI agent worse — a developer's case for write-side memory governance — Many_Audience7660 · 2026-08-24
- Gemini CLI Fix: Prevents Output Inflation on Negative maxChars — Kanika0306 · 2026-08-24
- Day 4 of Cloud Agents Migration: Tackling Constant PR Rebasing — jarrodwatts · 2026-08-24
- Graph Engineering organizes multi-agent systems via dynamic structures — Yuyuan Feng · 2026-08-24
- RecVerse agent simulates realistic e-commerce shopping sessions — Jiakai Tang · 2026-08-24