Agent Arena Analyzed 20,840 Coding-Agent Traces: 69% of Complaints Are Broken Code
arena · x · 2026-09-21
Agent Arena analyzed 20,840 traces from Agent Arena: Code across 29 models, using direct user praise and complaints to track the coding-agent frontier.
Key findings:
- Today's frontier models receive markedly more positive feedback; newer models across labs tend to show a more positive feedback balance.
- Broken code remains the biggest driver of complaints at 69.1%, followed by sloppy behavior. Other themes: incomplete output (27.1%), weak finish and usability (27.1%), and ignored instructions (20.3%).
- Models share common weaknesses but differ in how they disappoint: Fable 5.1 draws fewer "slop" and design complaints than Astra (3.9% vs 5.7% of sampled traces) and fewer "flaky" complaints (3.1% vs 5.0%).
The approach was inspired by the Claude Code team tracking profanities in user messages to gauge product experience. Author: Dawid Galarowicz.
More from coding & agent
- Zhejiang U & SJTU unveil DAS, an agent that writes publication-ready surveys in an hour — jiqizhixin · 2026-09-22
- Qwen team releases RecreationWorld: a five-platform sandbox for hybrid computer-use agents — _akhaliq · 2026-09-22
- Jev Skill Suggestion Injects Only the Skill Claude Code Needs, Saving Context — leslysandra · 2026-09-22
- Dev vibe-codes a living-room wall game with Codex, controlled by hand gestures — tristanbob · 2026-09-22
- A developer built an independent quality-and-popularity index for MCP servers — Prestigious-Web-2968 · 2026-09-22
- GitHub Copilot desktop app to get editable diffs for agent-generated changes — mariorod1 · 2026-09-22