Agent Arena leaderboard: Claude Fable 5.1 tops with +13.8% net improvement across 2M sessions
arena · x · 2026-09-26
Agent Arena's dynamic leaderboard ranks 44 models on real-world, long-horizon agent tasks using web, filesystem, and terminal tools, with 2M+ sessions collected. Top of the board: Claude Fable 5.1 (Max) at +13.80% net improvement ($4.11/task), GPT 6 Astra (Max) at +10.85% ($2.82), Claude Opus 5 (High) at +9.80% ($2.15), and GPT 6 Sol (Max) at +7.68% on just $0.74/task.
The board also tracks sub-signals including Confirmed Success, Steerability, Bash Recovery, and Tool Hallucination with confidence intervals.
More from coding & agent
- One agent writes the fix, another reviews it: a two-agent code review workflow in Slack — Al_Grigor · 2026-09-26
- Academic agent Memex upgraded to Opus 5.5: writing quality fixed, experience much better — arjunrajlab · 2026-09-26
- Dev uses open-source Ling-3.0-flash-VL to let AI redesign the foldable iPhone in a single HTML file — alifcoder · 2026-09-26
- Anthropic launches Claude plugin directory portal as MCP usage jumps 110x this year — ClaudeDevs · 2026-09-26
- Open-source Jev agent plays Pokemon Red live, pushing fast-decision AI beyond Tetris — supportingthedogs · 2026-09-26
- Anthropic deep dive: effort tuning in Claude Code pays off most for security and code review — trq212 · 2026-09-26