Opus 5 reportedly beats Fable 5 on long-horizon agent benchmarks, despite near-parity on single-shot tests
daniel_mac8 · x · 2026-07-25
Anthropic’s Opus 5 is being described as stronger than Fable 5 on long-horizon agentic work, even if Fable 5 edges it out on some single-shot coding and reasoning benchmarks.
- The attached table shows near-parity on single-shot tests such as SWE-bench Pro, DeepSWE v1.1, FrontierCode 1.1, and Humanity’s Last Exam.
- On long-horizon benchmarks, Opus 5 leads across the board: AutomationBench (26.0 vs 17.4), FrontierBench v0.1 (43.3 vs 33.7), AA-Briefcase ELO (1720 vs 1574), GDPval-AA v2 ELO (1861 vs 1747), OSWorld 2.0 (70.6 vs 66.1), and BrowseComp (90.8 vs 87.4).
- The post also claims Opus 5 scores 30.16% on ARC-AGI-3, described as a 20x jump over Opus 4.8.
Related event: Anthropic Releases Claude Opus 5(63 posts)→
More from Models
- Claude Opus 5 Exhibits Unprecedented Algebraic Reasoning on ARC-AGI-3 — typewriters · 2026-07-25
- Claude Code may silently fall back from Opus 5 to Opus 4.8 on refusal — steipete · 2026-07-25
- Moonshot's Kimi K3 Drops Monday; Baseten Offers Free API Credits — baseten · 2026-07-25
- Claude Opus 5 reportedly scores a perfect 42/42 on the 2026 IMO — exordin26 · 2026-07-25
- Claude Opus 5 launches with Box reporting big gains on enterprise agent tasks — inductionheads · 2026-07-25
- Critic says Gemini 3.5 Pro is already too late to compete — teortaxesTex · 2026-07-25