Coding Agent Index adds refusal views across DeepSWE, Terminal-Bench, SWE-Atlas-QnA
ArtificialAnlys · x · 2026-10-02
Companion post to Artificial Analysis's refusal-reporting launch, linking the full methodology: the Coding Agent Index v1.5 equally weights DeepSWE v1.1 (113 tasks), Terminal-Bench 4.0 (66 tasks), and SWE-Atlas-QnA (124 tasks), averaging pass@1 over three attempts per task across 31 models. Refusal timing and fallback views are available per benchmark, alongside time- and cost-per-task metrics.
Related event: Artificial Analysis Adds Safety Refusal Analysis to Coding Agent Index(2 posts)→
More from coding & agent
- Team of 8 AI agents beats best solo agent building Colosseum in Minecraft (0.72 vs 0.58) — DimitrisPapail · 2026-10-02
- AC2 adds automated reward-hacking monitors, AI agent investigates cheating RL runs — ypatil125 · 2026-10-02
- OpenAI admits Codex can't guarantee no subagent use, undermining result repeatability — RealSharpNinja · 2026-10-02
- Decagon launches Personal Agent Gateway and PACT protocol for AI-agent customers — kimberlywtan · 2026-10-02
- Why TDD works with AI coding agents: cheap guardrails and regression signals — randal_olson · 2026-10-02
- Economist runs full AI research pipeline in 45 minutes for just $19 with Expected Parrot agent — soumitrashukla9 · 2026-10-02