GPT-5.6 Sol fixes 18 planted regressions and 29 more in a 60-bug LMS benchmark

PawelHuryn · x · 2026-07-27

Repo 2: 60 shipped regressions, 7 frontier models

This second benchmark run used a different codebase: an LMS with 60 real regressions that had already shipped in production, each tied to the commit that originally fixed it.

Results from the screenshot and post:

The benchmark again uses strict diff grading, so the question is not whether the model sounded confident, but whether its edit actually fixed a real bug.

Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→

Original post →

More from Models

Models channel →