Small Models Write the Tests Too: 77% of Self-Written Test Suites Rejected Correct Code
deadatreides1 · reddit · 2026-10-05
A Reddit experiment with four tiny GGUF models (360M–1.7B, 785 calls, 6 tasks) found 129 of 168 self-written test suites (77%) rejected a known-correct reference solution, despite a 0.965 average mutation score — thorough but testing the wrong spec. Of 67 repair attempts, 50 patched already-correct code. Plain resampling (pass@8 = 0.958) beat the pipeline at similar token cost. The author's fix: judges come from human-written specs, not model-proposed tests.
Related event: Study Finds 77% of Model-Generated Tests Wrongly Reject Correct Code(2 posts)→
More from coding & agent
- NVIDIA Dynamo lets coding agents point at self-hosted endpoints with native tracing — TheZachMueller · 2026-10-06
- Are Companies Actually Deploying Business Agents Beyond Copilots? A Reddit Thread Asks for Production Truth — Most-Screen7458 · 2026-10-06
- Opus 5.5 One-Shots an Isometric Design in 5 Minutes, Skill Release Teased — KlausCodes · 2026-10-06
- Claude Appears to Have Added File Size Limits for In-Chat Downloads — nptacek · 2026-10-06
- marimo Studio: one notebook, separate audience views, agent-safe presentation layer — pandeyparul · 2026-10-06
- AI software factories: developers state intent, agents handle build and deploy — Pavan_Belagatti · 2026-10-06