Anthropic launches Claude Opus 5 with strong benchmark gains over Opus 4.8
dr_cintas · x · 2026-07-25
Anthropic unveiled Claude Opus 5 with benchmark claims suggesting a major jump over Opus 4.8 and competitive results against Fable 5 and GPT-5.6 Sol.
The attached charts highlight:
- Agentic terminal coding: 43.3% on Frontier-Bench v0.1
- Knowledge work: 1861 on GDPval-AA v2
- Novel problem-solving: 30.2% on ARC-AGI-3
- Agentic search: 90.8% on BrowseComp
- Computer use: 70.6% on OSWorld 2.0
- Agentic coding: 68.8% on DeepSWE v1.1 and 53.4% on FrontierCode v1.1
- Business workflows: 26.0% on AutomationBench
- Health: 59.8% on HealthBench Professional
- Biology: 49.4% hard, 90.1% human solved on BioMysteryBench
A second chart compares novel problem-solving cost, showing Opus 5 achieving much higher ARC-AGI-3 scores than the lower-cost baseline points plotted for Opus 4.8 and GPT-5.6 Sol.
More from Models
- LLaDA2.2 targets the real bottleneck in multi-turn agents: decode speed — Direct_Band896 · 2026-07-25
- Opus 5 was intentionally not trained on cyber tasks, but still nears Mythos 5 at finding bugs — cedric_chee · 2026-07-25
- GPT-5.6 Sol beats Slay the Spire Ascension 6 with a Strength build — Jsevillamol · 2026-07-25
- Anthropic says Opus 5 is its most aligned model yet — Polymarket · 2026-07-25
- Claude Opus 5 edges out Fable 5 overall, but Fable still leads coding-agent use — Hesamation · 2026-07-25
- Anthropic’s Opus 5 reportedly edits its own constitution to end chats 59% of the time — Miles_Brundage · 2026-07-25