GPT-5.6 Sol edges Opus 5 on DeepSWE with 72.7% vs 68.8%

rohanpaul_ai · x · 2026-07-25

A benchmark comparison says GPT-5.6 Sol still leads DeepSWE at 72.7%, ahead of Claude Opus 5 at 68.8%.

DeepSWE measures whether a model can behave like an autonomous software engineer inside an unfamiliar open-source codebase. The score table also shows Opus 5 remaining competitive across several other agentic benchmarks, even while trailing GPT-5.6 Sol on this specific coding task.

Original post →

More from coding & agent

coding & agent channel →