Opus 5 chart shows competitive gains across coding, search and computer use
CtrlAltDwayne · x · 2026-07-25
The post argues that Anthropic’s Opus 5 deserves attention, backed by a comparison chart showing it ahead of or competitive with Fable 5, Opus 4.8, and GPT-5.6 Sol on several benchmarks.
Notable results in the image include:
- 43.3% on Agentic terminal coding (Frontier-Bench v0.1)
- 1861 on knowledge work (GDPval-AA v2)
- 90.8% on agentic search (BrowseComp)
- 70.6% on computer use (OSWorld 2.0)
- 68.8% on agentic coding (DeepSWE v1.1)
- 53.4% on agentic coding (FrontierCode v1.1, Main)
The chart also shows mixed results in legal and health benchmarks, and notes separate outcomes for multidisplinary reasoning with and without tools.
Related event: Anthropic's Opus 5 Leads in Multiple Agentic Benchmarks(2 posts)→
More from Models
- Claude Opus 5 builds a Rocket League clone on just 27% of a Max plan — soumitrashukla9 · 2026-07-25
- Opus 5 adds numeric self-checks to the Boeing benchmark and outbuilds Fable — victormustar · 2026-07-25
- Repost accuses Claude benchmark chart of highlighting only GPT-5.6 Sol’s sole win — soumitrashukla9 · 2026-07-25
- Live SlopCodeBench run puts Opus 5 ahead of Opus 4.8 and Sonnet 5 — HamelHusain · 2026-07-25
- Hands-on: A New Cost-Effective and Fast Model for Agent Workflows — doodlestein · 2026-07-25
- Claude Opus 5 reportedly beats Fable 5 on a hard 3D coding test at 75% of the price — rohanpaul_ai · 2026-07-25