Debate Erupts Over Whether Agent Coding Evals Detect Regressions
Pawel Huryn argues that agent coding evals like DeepSWE are structurally flawed because agent-written tests are never re-run, so regressions go undetected. DeepSWE's author countered that tests serve beyond regression prevention and that agents introduced no extra regressions after legacy test suites were disabled.
2026-10-09 ~ 2026-10-09 · 3 related posts
- DeepSWE eval author: agents caused no more regressions even with existing test suites disabled — kunchenguid · 2026-10-09
- Pawel Huryn vs kunchenguid: do AI coding agents actually benefit from writing tests? — PawelHuryn · 2026-10-09
- PawelHuryn: DeepSWE's container-isolated tasks can't surface real-world regressions — PawelHuryn · 2026-10-09