Coding Agent Index adds refusal views across DeepSWE, Terminal-Bench, SWE-Atlas-QnA

ArtificialAnlys · x · 2026-10-02

Companion post to Artificial Analysis's refusal-reporting launch, linking the full methodology: the Coding Agent Index v1.5 equally weights DeepSWE v1.1 (113 tasks), Terminal-Bench 4.0 (66 tasks), and SWE-Atlas-QnA (124 tasks), averaging pass@1 over three attempts per task across 31 models. Refusal timing and fallback views are available per benchmark, alongside time- and cost-per-task metrics.

Related event: Artificial Analysis Adds Safety Refusal Analysis to Coding Agent Index(2 posts)→

Original post →

More from coding & agent

coding & agent channel →