Claude Opus 5 hits 74% on DeepSWE, topping long-horizon coding models
brandon_galang · x · 2026-07-29
The poster says Opus 5 is “objectively cracked” and should be studied by everyone using it, citing a quote that Claude Opus 5 has reached 74% on DeepSWE and is the best long-horizon coding model seen so far. The core claim is that Anthropic’s latest model is now leading this benchmark for long-horizon coding tasks.
More from Models
- Leak says GPT-6 slips to early September as Anthropic tests Fable 5.1 — soumitrashukla9 · 2026-07-29
- GPT-5.6 Sol Ultra finds a critical bug, then refuses to show it — haltakov · 2026-07-29
- Kimi K3 tops a benchmark chart in a repost claiming it beats Anthropic models — JarnoDuursma · 2026-07-29
- User Reports Grok's Generation Capabilities Have Gotten 'Real Cracked' — djcows · 2026-07-29
- User asks Anthropic not to deprecate Opus 4.6 until the model is fixed — oyacaro · 2026-07-29
- Hidden Trick: Manually Invoke Older Opus Models in Claude Code — voooooogel · 2026-07-29