Small Models Write the Tests Too: 77% of Self-Written Test Suites Rejected Correct Code

deadatreides1 · reddit · 2026-10-05

A Reddit experiment with four tiny GGUF models (360M–1.7B, 785 calls, 6 tasks) found 129 of 168 self-written test suites (77%) rejected a known-correct reference solution, despite a 0.965 average mutation score — thorough but testing the wrong spec. Of 67 repair attempts, 50 patched already-correct code. Plain resampling (pass@8 = 0.958) beat the pipeline at similar token cost. The author's fix: judges come from human-written specs, not model-proposed tests.

Related event: Study Finds 77% of Model-Generated Tests Wrongly Reject Correct Code(2 posts)→

Original post →

More from coding & agent

coding & agent channel →