DeepSWE Coding Agent Leaderboard: Claude Opus and GPT-5.6 Top the Charts
tristanbob · x · 2026-07-31
DeepSWE released a new leaderboard evaluating frontier coding agents on long-horizon engineering tasks across 113 tasks. Claude Opus and GPT-5.6-sol lead the board with pass rates of 74% and 73% respectively.
The leaderboard provides a detailed comparison of pass rate, average cost, output tokens, and agent steps:
- Claude Opus: 74% pass rate, avg cost $11.84
- GPT-5.6-sol: 73% pass rate, avg cost $8.39
- Claude Fable and GPT-5.6-terra tie for third (70%)
- Kimi-k3 follows closely with a 69% pass rate at a cost of $4.65
- GPT-5.6-luna achieves a 67% pass rate at an exceptionally low cost of $0.61
This benchmark offers a direct look at the capability limits and cost-effectiveness of current leading models in handling complex software engineering problems.
More from coding & agent
- Serving AI Agents Becomes a Storage and Networking Bottleneck, Starving GPUs — AccBalanced · 2026-07-31
- AUTOBOTS Announced: Recursively Self-Improving AI Workflows Without Human Intervention — bindureddy · 2026-07-31
- Aura Demo: 100% Local AI Agent Autonomously Controls Computers — ctjlewis · 2026-07-31
- NPort: An Open-Source Lightweight ngrok Alternative Powered by Cloudflare — tom_doerr · 2026-07-31
- DeepSeek-V4-Flash Officially Released, Surpassing Pro Preview in Agent Benchmarks — 赛博禅心 · 2026-07-31
- ComfyUI-QwenTTS Nodes Enable Voice Cloning and Design in ComfyUI — tom_doerr · 2026-07-31