Artificial Analysis opens up coding agent leaderboard and benchmarking methodology
ArtificialAnlys · x · 2026-09-19
Artificial Analysis published the full results page and methodology for its coding agent benchmarks. Coding Agent Index v1.5 combines three equally weighted benchmarks: DeepSWE v1.1 (113 software engineering tasks, by Datacurve), Terminal-Bench 4.0 (66 agentic terminal tasks, by Laude Institute), and SWE-Atlas-QnA (124 technical Q&A tasks, by Scale AI).
Each benchmark averages pass@1 across three attempts per task. The leaderboard also tracks time per task and API cost per task, and notes that agents with similar index values can differ significantly across repository tasks, terminal workflows, and rubric-based evaluations.
Related event: Coding Agent Index Adds Safety Refusal Metric; Claude Tops at 8.8%(2 posts)→
More from coding & agent
- macOS 27 ships with mlx_whisper built in, letting agents transcribe video locally — vista8 · 2026-09-19
- Custom /jev command in OpenCode turns LLM rambling into instant yes/no answers — TheMoonMidas · 2026-09-19
- Dev pipelines Codex output to Muse via Google Drive, calls it insane — RachelVT42 · 2026-09-19
- Building a Self-Improving Content Skill with Claude Code Artifacts and MCP — alexgoughcooper · 2026-09-19
- Jevable site catalogs 342 demos of Typeface's Jev, from $0.0039 flight search to intent-scored spreadsheets — gaganghotra_ · 2026-09-19
- Omar Khattab: AI-era programs still need real languages, declarative specs and DSPy-style modules — lateinteraction · 2026-09-19