GPT-5.6 Sol fixes 31 of 105 hidden bugs in a two-repo benchmark
PawelHuryn · x · 2026-07-27
What the benchmark measured
- The author ran 7 frontier models on 105 hidden bugs across two unrelated codebases: a VS Code extension and an LMS.
- Each model got one round per repo, with strict diff grading: a fix counted only if the bug was actually removed.
- The benchmark used ground-truth diffs, blind judges, anonymized submissions, and a withheld answer key.
Key results
- GPT-5.6 Sol: 31 bugs fixed total
- Fable 5: 24
- Opus 5: 21
- Kimi K3: 21
- Grok 4.5: 16
- Opus 4.8: 9
- Sonnet 5: 9
Notable takeaways
- In repo 1, 45 bugs were planted; 28 survived every model.
- In repo 2, 60 shipped regressions were used; 33 survived every model.
- The author argues that a green test suite does not prove a codebase is clean, and a model’s confident “no issues found” report can still miss many real bugs.
- The chart also shows large differences in wall-clock time and cost across models.
Related event: Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark(4 posts)→
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23