How Artificial Analysis builds its coding agent index: 3 benchmarks, cost and time
ArtificialAnlys · x · 2026-09-03
Artificial Analysis details the full data and methodology behind its coding agent leaderboard: Coding Agent Index v1.4 averages three benchmarks — DeepSWE (Datacurve, 113 SWE tasks), Terminal-Bench v2.1 (Laude Institute, 89 agentic terminal tasks), and SWE-Atlas-QnA (Scale AI, 124 technical Q&A tasks). Each benchmark averages pass@1 over three attempts per task, with equal weights. The leaderboard also shows average wall time and API cost per task, and the team cautions that agents with similar index values can still have very different strengths across repo tasks, terminal workflows, and rubric grading — read the per-eval breakdowns alongside the index.
More from Research
- Petar Veličković on Categorical Deep Learning: An Algebraic Theory of Architectures — burny_tech · 2026-09-03
- Petar Veličković on Categorical Deep Learning: An Algebraic Theory of Architectures — burny_tech · 2026-09-03
- NSF Renews Funding for Columbia's AI Earth-System Modeling Center LEAP for Five More Years — vishalmisra · 2026-09-03
- Cooperative AI publishes 'Priorities in Cooperative AI' talk by Lewis Hammond — ghadfield · 2026-09-03
- AirCaps Launches Audio Research Lab, Citing 20% WER for ASR on Noisy Real-World Speech — ycombinator · 2026-09-03
- SPAR to run RCTs testing whether secretly misaligned AI can sabotage human decisions — austinc3301 · 2026-09-03