Claude Opus 5 tops DeepSWE with a 74% score and a claimed 28% cost edge
daniel_mac8 · x · 2026-07-29
Claude Opus 5’s DeepSWE results are now out, and the model is being described as the top long-horizon coding agent on the benchmark.
- A shared chart shows Claude Opus 5 High near the top of DeepSWE, with an estimated 74% score.
- The comparison claims it beats GPT-5.6 Sol Max at roughly a 28% lower cost per task.
- The post argues that Opus 5 is now the most efficient option among the strongest long-task coding agents tested.
Related event: Claude Opus 5 Tops DeepSWE Benchmark(2 posts)→
More from coding & agent
- Inside OpenAI’s push to redesign software development for the agent era — danshipper · 2026-07-29
- opendot snapshots every shell action so terminal coding agents can be undone — alexriley12345 · 2026-07-29
- Open-source CLI agent goes public as more business-agent startups emerge — jasonkneen · 2026-07-29
- Recall speeds up RAG with episodic memory, cutting hits to 260ms from 49s — sharpeye_wnl · 2026-07-29
- Agent Wiki 0.8.0 adds multi-vault knowledge bases for agents — HockeyDadNinja · 2026-07-29
- Hamza masks secrets and PII before Claude Code or Codex sends data out — Suitable-Cow2000 · 2026-07-29