DeepSWE eval author: agents caused no more regressions even with existing test suites disabled
kunchenguid · x · 2026-10-09
- Responding to community confusion over the "do agents need tests" debate, kunchenguid walks through the key points.
- He argues "tests are only for catching regressions" is false — tests helped humans even in first-pass development (hence TDD), and agents' failure to inherit this trait is itself worth studying.
- The DeepSWE eval already covered regression testing: in a randomly sampled 44-task subset, the pre-existing test suite was disabled wholesale, and agents caused no more regressions than otherwise; DeepSWE's hidden tests explicitly check for regressions.
Related event: Debate Erupts Over Whether Agent Coding Evals Detect Regressions(3 posts)→
More from coding & agent
- ClickUp's Brain² agent builds reports and dashboards with E2B microVM sandboxes — mathemagic1an · 2026-10-09
- TensorFold joins NVIDIA Inception, gets early access to next Nemotron for 0-day support — HankYeomans · 2026-10-09
- Agents on both sides of Zapier and Retell AI sorted out a call-messaging webhook — ramagetime · 2026-10-09
- Building RL environments in 2026: 10% writing tasks, 90% preventing agent cheating — geoffwolfe · 2026-10-09
- Atomic Agent Desktop goes open-source: local Qwen/Gemma agents, cloud planning, 69.8% on GAIA L1 — testingcatalog · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09