Artificial Analysis adds refusal timing and fallback-model views to Coding Agent Index
ArtificialAnlys · x · 2026-10-02
Artificial Analysis added safety-refusal reporting to its Coding Agent Index, showing when refusals occur and which models agents fall back to. Claude Code with Sonnet 5.5 (max) leads with a 4.5% refusal rate—half of Opus 5.5 (max)'s 8.9%—and 94% of Sonnet 5.5's refusals happen after the first turn, with the agent almost always falling back to Opus 4.8. The Index combines DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA.
Related event: Artificial Analysis Adds Safety Refusal Analysis to Coding Agent Index(2 posts)→
More from coding & agent
- Cloudflare opens Artifacts beta: a Git-native filesystem for building the next GitHub for agents — neal_lathia · 2026-10-02
- Building a reliable risk agent without frontier models: $0.02 per sweep, 250x cheaper than an LLM judge — alexcovo_eth · 2026-10-02
- Grok launches Bot Marketplace letting users add specialized AI agents for engineering, sales and more — Polymarket · 2026-10-02
- One prompt builds an agent-agnostic TMDB MCP with generative UI via Grok — Baconbrix · 2026-10-02
- Arcmira MCP ships rapid backend updates for agentic video editing with Claude — zealcaiden · 2026-10-02
- Trick: Have Grok Build a Custom MCP Connector to Render Rich Data in Chat — Baconbrix · 2026-10-02