Agent Arena Launches Agent Leaderboard Ranking 46 Models on 2M+ Real-World Sessions
arena · x · 2026-09-29
Agent Arena has launched a performance leaderboard that measures how well AI models orchestrate tools on real-world, long-horizon tasks using web, filesystem, and terminal tools, based on over 2 million sessions across 46 models.
Key takeaways
- Claude Fable 5.1 (Max) tops the board with 14.06% net improvement and 17.37% confirmed success rate
- Claude Opus 5.5 (High) ranks second (11.84%) at a notably low median cost of $1.33 per task
- GPT 6 Astra (Max) takes third (10.36%) with the highest praise-vs-complaint ratio (33.92%)
- Anthropic models dominate the top 10; per-task costs range roughly from $0.81 to $3.47
Metrics include tool reliability, steerability, bash recovery, and tool hallucination rate, making it a practical reference for choosing agent models.
More from Models
- Opus 5.5 tops Drone-Bench and cheats far less than prior Claude models — scaling01 · 2026-09-29
- User claims 'Opus 5.5' turned a post on agent harnesses into an explainer video in one shot — alex_verem · 2026-09-29
- Engineer proud as Sonnet 5.5 scores 61.6% on chartography benchmark — echen · 2026-09-29
- ProgramBench multi-agent eval: Opus 5.5 fastest with a 5-agent team, Sonnet 5.5 with subagents — jyangballin · 2026-09-29
- Arrow 2 Telos tops Design Arena's SVG generation benchmark — AWizardWhoCodes · 2026-09-29
- Emulate-1 claims to beat AI detectors: outputs pass Pangram as human writing — alejandroll10 · 2026-09-29