Hidden two-file test: 3 of 8 agent fixes closed the reported bug by breaking another

lulzxdxdxd · reddit · 2026-10-10

The author ran seeded-bug experiments across 4 models (2 trials each, temperature 0) with a hidden test suite spanning two files: all 8 runs closed the reported failure, but 3 of 8 did so by breaking a different test.

The failure mode: the model patches the exact line the traceback names, the symptom disappears — while the test it broke only fails once the first bug is gone. A single-file task can't reveal this; a two-file task can.

The checklist he now uses:

The striking result: rankings on the reported failure differed from rankings on the full suite, and only the latter predicted which model was still in the codebase three weeks later. If your eval reports only the failure you asked about, the model you keep won't be the one your eval ranked first.

Original post →

More from coding & agent

coding & agent channel →