GPT-5.6 Sol fixes 18 planted regressions and 29 more in a 60-bug LMS benchmark
PawelHuryn · x · 2026-07-27
Repo 2: 60 shipped regressions, 7 frontier models
This second benchmark run used a different codebase: an LMS with 60 real regressions that had already shipped in production, each tied to the commit that originally fixed it.
Results from the screenshot and post:
- GPT-5.6 Sol fixed 18 planted regressions, and then 29 more bugs that had not been planted by the benchmark author.
- Other models landed between 1 and 6 fixes in the author’s summary, while the chart in the image shows the rest of the per-model breakdown.
- The author says they thought the codebase was clean before running the models.
The benchmark again uses strict diff grading, so the question is not whether the model sounded confident, but whether its edit actually fixed a real bug.
Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27