Opus 5 posts 30.2% on ARC-AGI-3 and leads several agentic benchmarks
socoolandawesome · reddit · 2026-07-25
Opus 5 posts a strong benchmark run
The image attached to the post shows Opus 5 ahead on several benchmarks, including:
- 43.3% on Agentic Terminal Coding (Frontier-Bench v0.1)
- 1861 on Knowledge Work (GDPval-AA v2)
- 30.2% on ARC-AGI-3
- 90.8% on Agentic Search (BrowseComp)
- 64.7% with tools on Humanity’s Last Exam
- 70.6% on OSWorld 2.0
- 68.8% on DeepSWE v1.1
- 53.4% on FrontierCode v1.1 Main
The chart compares Opus 5 against Fable 5, Opus 4.8, and GPT-5.6 Sol, presenting Opus 5 as stronger on most listed tasks, especially agentic coding, search, and novel problem solving.
Related event: Anthropic Releases Claude Opus 5(40 posts)→
More from Models
- Claude Opus 5 lands on AWS Bedrock with ZDR and production APIs — AWS ML Blog · 2026-07-25
- Opus 5 is shown as a new Pareto-optimal LLM with strong ARC-AGI-3 results — brandon_galang · 2026-07-25
- Anthropic says Claude Opus 5 matches frontier intelligence at half the price — burny_tech · 2026-07-25
- Opus 5 is being compared to Opus 4.8 with a two-month gap and big benchmark jumps — SuhailKakar · 2026-07-25
- Claude Opus 5 lands at half the price and becomes the default on Claude Max — minchoi · 2026-07-25
- Claude Opus 5 reportedly kept “crystallizing” for 27 minutes in an initial test — Daniel_Farinax · 2026-07-25