Bug Hunt Bench: blind-graded leaderboard testing frontier coding models on 105 real planted bugs
PawelHuryn · x · 2026-10-08
The author launched Bug Hunt Bench, a persistent leaderboard for frontier coding models tested on 105 real planted bugs. Highlights: blind grading, one prompt per repo, ranked by bugs fixed; sortable by cost, time, turns, and coverage; supports multiple reasoning tiers and harnesses, with superseded historical runs excluded from the default view. Full run notes and data are public on GitHub (run-notes.md, runs.csv). Built by Pawel Huryn of The Product Compass newsletter.
More from Models
- Dev finds 6.1 Sol surprisingly good at designing native iOS apps — Dimillian · 2026-10-08
- Claude Code lead resets usage limits after community vote: users pick capacity over new features — CurieuxExplorer · 2026-10-08
- Prompting Opus 5.5 with a "grumpy senior engineer reviewer" kept it benchmarking for 2 days — remilouf · 2026-10-08
- Claude Haiku 5.5 beats GPT-6 Luna at matching price, but burns ~3x the tokens — Latent Space · 2026-10-08
- Gemini 3.8 Flash: 3 images and 5 chat messages already eat 8% of the quota — Horror-Airport-7606 · 2026-10-08
- Haiku 5.5 benchmarks 2-3x faster than GPT-6 Luna and performs better — PawelHuryn · 2026-10-08