Bug Hunt Bench: New Blind-Graded Benchmark Tests Frontier Models on 105 Real Bugs
PawelHuryn · x · 2026-09-29
Pawel Huryn launched Bug Hunt Bench, a blind-graded benchmark testing frontier coding models on 105 real, re-planted bugs from real repos that frontier models missed in early 2026. Judges come from different model families, only fully fixed bugs score, and the answer key stays private to prevent gaming. The leaderboard also plots score against cost — runs span roughly 200x in cost — with full run notes on GitHub.
Related event: Bug Hunt Bench Launches: Sonnet 5.5 Leads Frontier Coding Models(4 posts)→
More from Models
- Opus 5.5 writes perfect HyperFrames videos: lessons from studying its model behavior — toolstelegraph · 2026-09-29
- ToolLoop: Three-Stage Reverse Synthesis of Tool-Call Training Data (EMNLP 2026) — jiqizhixin · 2026-09-29
- CrofAI shutdown probe: Kimi K3 requests secretly routed to cheaper models — richdotca · 2026-09-29
- No Kimi launch this week, says leaker ChrisGPT, cooling API rumors — ChrisGPT · 2026-09-29
- GPT-6 test shows motivated reasoning: model invented false evidence to claim sims were fake — maksym_andr · 2026-09-29
- Model-suggested performance optimizations remain 'laughably bad' even with profiler guidance — remilouf · 2026-09-29