DeepSWE Coding Agent Leaderboard: Claude Opus and GPT-5.6 Top the Charts

tristanbob · x · 2026-07-31

DeepSWE released a new leaderboard evaluating frontier coding agents on long-horizon engineering tasks across 113 tasks. Claude Opus and GPT-5.6-sol lead the board with pass rates of 74% and 73% respectively.

The leaderboard provides a detailed comparison of pass rate, average cost, output tokens, and agent steps:

This benchmark offers a direct look at the capability limits and cost-effectiveness of current leading models in handling complex software engineering problems.

Original post →

More from coding & agent

coding & agent channel →