Claude Opus 5 tops AA-Briefcase with 1720 Elo and 20% lower task cost than Fable 5
ArtificialAnlys · x · 2026-07-25
Artificial Analysis says Claude Opus 5 is the new leader on its AA-Briefcase agentic knowledge-work benchmark, beating Claude Fable 5 by about 146 Elo at max effort.
Key results
- Max effort: 1720 Elo on AA-Briefcase, versus 1574 for Claude Fable 5
- Cost per task: $17.79, about 20% cheaper than Fable 5’s $22.30
- High-effort mode: $10.41 per task while still beating Fable 5 by 32 Elo
- Analytical quality: 2016 Elo, nearly 300 points ahead of Fable 5
- Presentation quality: 1628 Elo, still about 40 points behind GPT-5.6 Sol (max)
The benchmark is based on realistic private knowledge-work tasks with thousands of files and deliverables like reports, slides, and spreadsheets. Artificial Analysis says Opus 5’s gains come mainly from rubric pass rate and analytical quality, while the tradeoff is slower runtime and more turns per task.
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11