Seven frontier models miss 44 hidden bugs even with 922 green tests
PawelHuryn · x · 2026-07-27
Repo 1: 45 hidden bugs, 7 models, 1 round each
The benchmark tested seven frontier models on a VS Code extension with 45 hidden bugs. All 922 tests stayed green throughout, but the diff-based grading showed that a green test suite did not mean the code was actually fixed.
Key takeaways from the run:
- Sonnet 5 fixed just 1 bug, added a regression test for it, and then reported that nothing else looked suspicious.
- 44 bugs were still left untouched.
- The screenshot shows the per-model fixes, wall-clock time, and cost, with GPT-5.6 Sol leading the field on this repo.
The post’s point is blunt: a passing test suite and a confident model self-report are both insufficient as proof of correctness.
Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23