Opus 5 arrives with stronger health answers and a benchmark chart spanning coding, search, and biology
JeremyNguyenPhD · x · 2026-07-25
The post says Opus 5 is out and that, unlike Fable, it can answer medical questions.
The attached benchmark chart compares Opus 5 with Fable 5, Opus 4.8, and GPT-5.6 Sol across several tasks:
- Agentic terminal coding: 43.3%
- Knowledge work (GDPval-AA v2): 1861
- Novel problem-solving (ARC-AGI-3): 30.2%
- Agentic search: 90.8%
- Multidisciplinary reasoning: 56.3% without tools, 64.7% with tools
- Computer use: 70.6%
- Agentic coding: 68.8% on DeepSWE v1.1, 53.4% on FrontierCode v1.1 Main
- Business workflows: 26.0%
- Legal: 11.7%
- Health: 59.8%
- Biology: 49.4% hard, 90.1% human solved
The visual positions Opus 5 as especially strong on agentic search, computer use, and several reasoning benchmarks, while also highlighting a health-related capability that the author says is better than Fable’s.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(97 posts)→
More from Models
- Claude Opus 5 is shown hitting 449.46× speedup on a kernel benchmark — scaling01 · 2026-07-25
- Hamel Husain says evals beat vibes after Opus 5 costs 6× more and scores worse — HamelHusain · 2026-07-25
- Claim says Claude Opus 5 scored 42/42 on the 2026 International Math Olympiad — Polymarket · 2026-07-25
- Claude Opus 5 feels better at turning messy inputs into a sharp insight, user says — iruletheworldmo · 2026-07-25
- Claude Opus 5 gets treated like it has “taste” in a meme-style repost — repligate · 2026-07-25
- Claude Opus 5 is being shared for a quirky “back to work” behavior clip — repligate · 2026-07-25