Bug Hunt Bench: GPT-5.6 Sol fixes 22 bugs, Opus 5 12, on a 45-bug repo
PawelHuryn · x · 2026-07-25
A Bug Hunt Bench run on a VS Code extension with 45 hidden bugs compares six coding models under the same prompt and setup.
- GPT-5.6 Sol leads with 13 + 9 fixes, taking 21m57s and costing $10.76.
- Opus 5 fixes 11 + 1, in 17m22s for $15.89.
- Fable 5 gets 9 + 2 but is the most expensive at $36.42.
- Grok 4.5 fixes 5 + 2 for $5.82.
- Kimi K3 lands at 4 + 0 after 62m19s.
- Opus 4.8 fixes 2 + 0.
The test setup kept the 922-test suite green, and 27 of the 45 bugs survived all six models, showing the task is still hard even for the best systems.
Related event: GPT-5.6 Beats Claude Opus 5 in Bug Hunt Benchmark(3 posts)→
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11