AutomationBench chart puts Opus 5 ahead on long-horizon agent tasks

daniel_mac8 · x · 2026-07-25

The post argues that “medium” thinking effort should be the default for this model, and the attached benchmark image compares agentic business workflow performance on AutomationBench.

The chart places Opus 5 above the other listed models on long-horizon business workflow tasks, with pass rates roughly in the 22–26% range depending on effort settings. The reply adds a comparison claim: Fable 5 may be slightly better at single-shot tasks, but Opus 5 is stronger on agentic benchmarks that measure long-term task completion.

Related event: Anthropic Releases Claude Opus 5(63 posts)→

Original post →

More from Models

Models channel →