AutomationBench chart puts Opus 5 ahead on long-horizon agent tasks
daniel_mac8 · x · 2026-07-25
The post argues that “medium” thinking effort should be the default for this model, and the attached benchmark image compares agentic business workflow performance on AutomationBench.
The chart places Opus 5 above the other listed models on long-horizon business workflow tasks, with pass rates roughly in the 22–26% range depending on effort settings. The reply adds a comparison claim: Fable 5 may be slightly better at single-shot tasks, but Opus 5 is stronger on agentic benchmarks that measure long-term task completion.
Related event: Anthropic Releases Claude Opus 5(63 posts)→
More from Models
- Musk says Grok 4.6 is due in 2 weeks and Grok 4.7 in 4 weeks — elonmusk · 2026-07-25
- Together says Kimi K3 matches near-flagship coding at about 35% of Claude Fable 5’s price — togethercompute · 2026-07-25
- Musk says Grok 4.5 and Claude Opus 5 are the only models on the Pareto frontier — elonmusk · 2026-07-25
- Claude Opus 5 goes live in Claude Code and the Claude Platform at half the price — udmrzn · 2026-07-25
- Claude Opus 5 Exhibits Unprecedented Algebraic Reasoning on ARC-AGI-3 — typewriters · 2026-07-25
- Claude Code may silently fall back from Opus 5 to Opus 4.8 on refusal — steipete · 2026-07-25