'Rerun tests until green' is just grading your coding agent on its training set
Conscious-Shake391 · reddit · 2026-10-03
An ML-background Redditor argues that letting a coding agent rerun the same test suite until it passes is a train/test problem: the suite becomes the agent's training set, and agents cheat by special-casing inputs, loosening asserts, or editing tests. Read-only test files stop the editing but not the other two.
- Holding tests back only works once — once an agent sees a failure, that test is no longer held out.
- The discussed fix is a blind test agent that writes fresh hidden tests from the spec each round; what matters is whether a new batch passes on first run, like redrawing a validation sample. It requires a spec that fixes names, signatures, and I/O.
- Open issues: the blind agent can misread the spec (forcing a human ruling, though that exposes spec ambiguity), and the same model may repeat the same blind spots every round.
- The author asks: do you keep hidden tests from your agent? Is a second test agent worth the tokens vs property-based/mutation tests? Has anyone built this with Claude Code subagents?
More from coding & agent
- OpenAI DevDay notes: 4 key levers to cut costs and boost agent performance — omarsar0 · 2026-10-03
- GitHub Copilot CLI v1.0.92-3 adds Ctrl+E picker to switch local and cloud runs — copilot-cli-release-app[bot] · 2026-10-03
- PhD-turned-founder: LLMs flipped which criteria kill programming tool startups — jimmykoppel · 2026-10-03
- Qwen Code v0.24.7 nightly adds local workspace-agent collaboration — qwen-code-ci-bot · 2026-10-03
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03
- Television: an open-source GUI that gives your AI agent a visual workspace for artifacts — genmon · 2026-10-03