Model-written tests rejected a known-correct solution 77% of the time in agent pipeline test

deadatreides1 · reddit · 2026-10-05

A developer built the textbook coding-agent pipeline (contract → tests → code → run tests → repair), then fed every generated test suite a known-correct reference solution: 129 of 168 suites (77%) rejected it.

Key findings:

Caveats: tiny local models (360M-1.7B), 6 tasks. The author's takeaway: the deciding test must come from the spec or human-written examples — a model may propose tests, but it doesn't get to be the judge.

Original post →

More from coding & agent

coding & agent channel →