Agent Arena leaderboard: 2M+ real-world agentic tasks rank Claude Fable 5.1 on top
arena · x · 2026-09-29
Agent Arena launched a leaderboard measuring models on millions of real-world, long-horizon agentic tasks — 2,058,478 sessions across 45 models — where models use web search, filesystem, and terminal tools, scored via causal tracing of outcomes relative to the average model.
Top of the board (Net Improvement):
- Claude Fable 5.1 (Max) 13.84%, 17.51% confirmed success, $3.51 median cost/task
- Claude Opus 5.5 (High) 12.15%, best steerability (14.5%), $1.35/task
- GPT 6 Astra (Max) 10.31%
- Claude Opus 5 (Max) 9.58%
- Claude Opus 5 (High) 9.47%
Sub-metrics include tool reliability, praise vs complaint, Bash recovery, and tool hallucination (all labs around 0.3%), with Anthropic models leading on confirmed success rates.
More from Models
- JevBench v1.5: Cygnet and Winnow-12B tie for the top spot in a first for the leaderboard — airesearch12 · 2026-09-29
- DiffusionGemma becomes an open, self-hostable decision model on vLLM amid Jev's viral 'System One' moment — rseroter · 2026-09-29
- Open source doubles M5 Ultra MLX token prefill in just one week via Flash-Next — TheMoonMidas · 2026-09-29
- All three OpenAI GPT-6 models — Astra, Sol, Luna — go live on Runware's endpoint — aziz4ai · 2026-09-29
- Opus 5.5 one-shots 3D explainers, seen as finally good enough for an 'Diamond Age' autotutor — anselm · 2026-09-29
- Open-source Gemma 4 voice translator runs fully offline on a Raspberry Pi — tom_doerr · 2026-09-29