Model-written tests rejected a known-correct solution 77% of the time in agent pipeline test
deadatreides1 · reddit · 2026-10-05
A developer built the textbook coding-agent pipeline (contract → tests → code → run tests → repair), then fed every generated test suite a known-correct reference solution: 129 of 168 suites (77%) rejected it.
Key findings:
- The tests weren't lazy: average mutation score 0.965, and 81% of suites with a perfect 1.0 mutation score still failed the correct answer — thorough, confident, and testing the wrong spec.
- Favorite case: on a vowel-counting task, the tests and the bug shared the same misunderstanding (both expected 5 instead of 10 for mixed-case input), so the buggy code PASSED and repair never ran.
- Of 67 repair attempts, 50 were fixing code the tests had falsely rejected; only 1 of 17 real bugs was actually fixed.
- Better models false-reject less but still badly: llama-3.2-1b 0.92-1.0, qwen2.5-coder-1.5b 0.75, qwen3-1.7b 0.47-0.65.
- The whole pipeline didn't beat a dumb 4-8x sampling baseline at similar token budget.
Caveats: tiny local models (360M-1.7B), 6 tasks. The author's takeaway: the deciding test must come from the spec or human-written examples — a model may propose tests, but it doesn't get to be the judge.
More from coding & agent
- Redditor drafts yt-dlp + Whisper + SQLite pipeline to auto-summarize expert content — QuestionAsker2030 · 2026-10-05
- SpaceX engineer: AI agents now merge PRs on their own — over 1,000 a month — itsOmSarraf_ · 2026-10-05
- Astra agent reverse-engineers macOS binaries to clone all 27 Liquid Glass effects in GPUI — steipete · 2026-10-05
- Local MCP server gives coding agents a project map: 9.3x fewer tokens on a real 12-step task — Honest_Traffic_8613 · 2026-10-05
- Jeff Dean's 1-hour AI engineering lecture: from LLM basics to graph-orchestrated agent swarms — glenbeer · 2026-10-05
- Fully local parkour sim vibe-coded with GLM 5.3 Flash on 2x DGX Sparks, recipe open-sourced — -dysangel- · 2026-10-05