GPT-5.6 Sol fixes 18 planted regressions and 29 more in a 60-bug LMS benchmark
PawelHuryn · x · 2026-07-27
Repo 2: 60 shipped regressions, 7 frontier models
This second benchmark run used a different codebase: an LMS with 60 real regressions that had already shipped in production, each tied to the commit that originally fixed it.
Results from the screenshot and post:
- GPT-5.6 Sol fixed 18 planted regressions, and then 29 more bugs that had not been planted by the benchmark author.
- Other models landed between 1 and 6 fixes in the author’s summary, while the chart in the image shows the rest of the per-model breakdown.
- The author says they thought the codebase was clean before running the models.
The benchmark again uses strict diff grading, so the question is not whether the model sounded confident, but whether its edit actually fixed a real bug.
Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23