Bug Hunt Benchmark: GPT-5.6 and Opus 5 Lead in Code Repair
A Bug Hunt Bench test on 45 hidden bugs in a VS Code extension repository revealed significant performance differences among AI models. GPT-5.6 Sol successfully fixed 22 bugs, while Claude Opus 5 fixed 11 to 12, and Opus 4.8 managed only 2.
2026-07-25 ~ 2026-07-25 · 2 related posts
- Bug Hunt Bench: GPT-5.6 Sol fixes 22 bugs, Opus 5 12, on a 45-bug repo — PawelHuryn · 2026-07-25
- Claude Opus 5 fixes 11 of 45 hidden bugs, versus 2 for Opus 4.8 — PawelHuryn · 2026-07-25