EvoCode-Bench: Multi-turn Coding Pass Rates Plunge to 7.7% by Round 10
dl_weekly · x · 2026-08-03
While traditional coding benchmarks like SWE-bench rely on single-turn prompts, EvoCode-Bench is designed to evaluate AI agents in multi-turn interactions. The benchmark consists of 26 tasks spanning 227 sequential rounds, requiring agents to handle evolving instructions (such as extending codebases, correcting logic, or managing conflicting requirements) within a persistent workspace, followed by cumulative testing after each turn.
Empirical results show a dramatic drop in pass rates, from 46.7% in Round 1 to just 7.7% by Round 10. Interestingly, failures are primarily driven by regressions rather than missing features, exposing the fragility of current AI coding agents in long-term, iterative workflows.
More from coding & agent
- Developer Shares AI Coding Workflow: Shipping Iterations Without Reading Code — dfinke · 2026-08-03
- DeepSeek V4 Flash Performance Varies Wildly Across Coding Agents — PMinervini · 2026-08-03
- Fixing Coding Agents' Visual Blind Spots: Introducing SceneProof — ReyJ94 · 2026-08-03
- Fully Prompt-Driven: Dev Builds Complete 3D RPG Using AI-Generated Three.js Code — majidmanzarpour · 2026-08-03
- Dev Says If AI Only Resolved Merge Conflicts, It'd Still Be Worth It — cantrell · 2026-08-03
- METATRON: Open-Source Penetration Testing Agent Powered by Local LLMs — tom_doerr · 2026-08-03