Pawel Huryn vs kunchenguid: do AI coding agents actually benefit from writing tests?
PawelHuryn · x · 2026-10-09
- Pawel Huryn argues current coding-agent evals are structurally limited: the agent's tests get exactly one shot, the grader never runs them, nothing comes after — so the eval can't tell whether agents would catch the next regression.
- kunchenguid counters that regression was already covered: when the pre-existing test suite was disabled wholesale, the agent caused no more regressions, and deepswe's hidden tests explicitly check for regressions.
- Huryn holds his conclusion: the eval only proves agents can make a single change without tests (they understand code without compiling it), saying nothing about tests' long-term value in maintenance.
Related event: Debate Erupts Over Whether Agent Coding Evals Detect Regressions(3 posts)→
More from coding & agent
- Inherit-MAS cuts multi-agent token use by up to 34.6% with evolution-inspired inheritance — Songtao Wei · 2026-10-09
- Agent societies in empty 3D worlds: DeepSeek writes a charter, Gemini builds a bridge — pkmital · 2026-10-09
- Local Hermes agent scans whole system, finds 90% of bookkeeping docs in 5 minutes — natesiggard · 2026-10-09
- Why Shipping Hundreds of PRs Makes Sense: AI Removes the Intelligence Bottleneck — vinvan · 2026-10-09
- A 9-step dependency plan, one voice prompt, 8 parallel agent threads spawned — Yamapama · 2026-10-09
- Every shares its 4-step Claude + Hyperframes workflow for turning articles into social videos — every · 2026-10-09