Hidden two-file test: 3 of 8 agent fixes closed the reported bug by breaking another
lulzxdxdxd · reddit · 2026-10-10
The author ran seeded-bug experiments across 4 models (2 trials each, temperature 0) with a hidden test suite spanning two files: all 8 runs closed the reported failure, but 3 of 8 did so by breaking a different test.
The failure mode: the model patches the exact line the traceback names, the symptom disappears — while the test it broke only fails once the first bug is gone. A single-file task can't reveal this; a two-file task can.
The checklist he now uses:
- Does the patch only touch the file the traceback named, when the cascade lives in the second file?
- Run the whole suite after every patch, not once at the end — aggregate pass/fail hides "passes its own test, breaks a neighbour" patches
- Score attempts and outcomes separately: two tries ending green beats one "first-try fix" leaving a red neighbour
The striking result: rankings on the reported failure differed from rankings on the full suite, and only the latter predicted which model was still in the codebase three weeks later. If your eval reports only the failure you asked about, the model you keep won't be the one your eval ranked first.
More from coding & agent
- Bend's author: a ~100k-token kernel designed so AI agents can rewrite the compiler — MikePFrank · 2026-10-10
- Gave Claude $1,000 and my therapy transcripts to plan my birthday — tech__unicorn · 2026-10-10
- Worried about agents wiping your Mac? Hourly Time Machine backups, says dev — HankYeomans · 2026-10-10
- Creator uses Claude Code with ElevenLabs and GPT Image 2 to auto-generate 8 cartoon clips in 8 art styles — mhmazur · 2026-10-10
- Claude Loop Engineering: four loop types and the four bills nobody warns you about — blaizedsouza · 2026-10-10
- Diablo 2 Resurrected iOS port finished, GitHub repo coming in days — sujingshen · 2026-10-10