Coding Agent Leaderboard Updated
WenhuChen · x · 2026-07-09
Artificial Analysis has updated its Coding Agent Index, replacing the previous SWE-Bench Pro with Datacurve's DeepSWE. Following the update, GPT-5.5's Codex surpassed Claude Code with Opus 4.8 on the leaderboard, while the newly released Claude Fable 5 also debuted at the top within Claude Code.
The post further explains the reason for the swap: DeepSWE generates tasks from scratch rather than reusing public GitHub issues or PRs, making it much harder for models to simply "memorize the answers." In contrast, SWE-Bench Pro had become increasingly vulnerable to "cheating," as models could exploit the repository's commit history to solve tasks.
Related event: OpenAI Launches GPT-5.6 Series: Multi-Agent and Cost-Efficiency(119 posts)→
More from coding & agent
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11
- hyperresearch: agent-driven knowledge base that turns web research into a searchable wiki — jordan-gibbs · 2026-09-11
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11