Claude Opus 5 posts 96.0% on SWE-bench Verified in new score chart
connoraxiotes · x · 2026-07-25
- The post references a chart from Claude showing Opus 5 as highly efficient on SWE-bench.
- In the image, Opus 5 is reported at 96.0% on SWE-bench Verified, 79.2% on SWE-bench Pro, 89.5% on SWE-bench Multilingual, and 59.4% on SWE-bench Multimodal.
- The joke is that the benchmark graph feels decisive enough to “rest easy” on SWE-bench Verified.
- This is a model-performance post, with the benchmark numbers doing the real work.
Related event: Anthropic Releases Claude Opus 5(40 posts)→
More from Models
- Anthropic’s Claude releases appear to have sped up from every four months to monthly in 2026 — dustinvtran · 2026-07-25
- Claude Opus 5 lands on AWS Bedrock with ZDR and production APIs — AWS ML Blog · 2026-07-25
- Opus 5 is shown as a new Pareto-optimal LLM with strong ARC-AGI-3 results — brandon_galang · 2026-07-25
- Anthropic says Claude Opus 5 matches frontier intelligence at half the price — burny_tech · 2026-07-25
- Opus 5 is being compared to Opus 4.8 with a two-month gap and big benchmark jumps — SuhailKakar · 2026-07-25
- Claude Opus 5 lands at half the price and becomes the default on Claude Max — minchoi · 2026-07-25