Artificial Analysis opens up coding agent leaderboard and benchmarking methodology

ArtificialAnlys · x · 2026-09-19

Artificial Analysis published the full results page and methodology for its coding agent benchmarks. Coding Agent Index v1.5 combines three equally weighted benchmarks: DeepSWE v1.1 (113 software engineering tasks, by Datacurve), Terminal-Bench 4.0 (66 agentic terminal tasks, by Laude Institute), and SWE-Atlas-QnA (124 technical Q&A tasks, by Scale AI).

Each benchmark averages pass@1 across three attempts per task. The leaderboard also tracks time per task and API cost per task, and notes that agents with similar index values can differ significantly across repository tasks, terminal workflows, and rubric-based evaluations.

Related event: Coding Agent Index Adds Safety Refusal Metric; Claude Tops at 8.8%(2 posts)→

Original post →

More from coding & agent

coding & agent channel →