Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark

A benchmark of seven frontier models across two codebases revealed significant limitations in AI debugging, with 63 out of 105 hidden bugs remaining undetected by any model. GPT-5.6 Sol performed best but still managed to fix only 31 bugs.

2026-07-27 ~ 2026-07-27 · 4 related posts