EvoCode-Bench: Multi-turn Coding Pass Rates Plunge to 7.7% by Round 10

dl_weekly · x · 2026-08-03

While traditional coding benchmarks like SWE-bench rely on single-turn prompts, EvoCode-Bench is designed to evaluate AI agents in multi-turn interactions. The benchmark consists of 26 tasks spanning 227 sequential rounds, requiring agents to handle evolving instructions (such as extending codebases, correcting logic, or managing conflicting requirements) within a persistent workspace, followed by cumulative testing after each turn.

Empirical results show a dramatic drop in pass rates, from 46.7% in Round 1 to just 7.7% by Round 10. Interestingly, failures are primarily driven by regressions rather than missing features, exposing the fragility of current AI coding agents in long-term, iterative workflows.

Original post →

More from coding & agent

coding & agent channel →