Across 105 hidden bugs, 63 survived every frontier model
PawelHuryn · x · 2026-07-27
63 of 105 bugs survived every model
This follow-up adds the benchmark’s methodology and the aggregate result across the two repos:
- 105 hidden bugs were tested in total.
- 63 bugs survived every model.
- The benchmark used blind judges, anonymized submissions, and a withheld answer key.
- The author stresses that the diff is the ground truth, not the model’s own self-report.
- 43 of the 63 surviving bugs had already shipped in a real product and had been fixed once before.
The takeaway is that multiple frontier models can reread the same code and still miss bugs that are already known in production history.
Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23