Bug Hunt Bench ranks frontier coding models on 105 real planted bugs, with 200x cost spread
PawelHuryn · x · 2026-09-10
- Pawel Huryn, author of The Product Compass, runs Bug Hunt Bench: a blind-graded benchmark that plants real bugs in real repos and ranks frontier coding models by how many of the 105 they fix, with one prompt per repo.
- The leaderboard offers score-vs-cost, score-vs-time, and coverage views; cost is on a log axis because the spread across models is roughly 200x, while runtime spans under 7x.
- In replies, Huryn clarifies a deepseek-v4-pro run was already tested (just filtered out of the default view) and another is queued to publish automatically after two runs finish.
More from Models
- DeepSeek V4.1 paper praised as a top-5 DeepSeek paper, textbook-style — teortaxesTex · 2026-09-10
- Google AI Mode now cites 72% fewer sources for logged-out users, data shows — gaganghotra_ · 2026-09-10
- DeepSeek cites 2024 YoCo paper as inspiration behind its CED transformer blocks — jm_alexia · 2026-09-10
- Research has cut LLM costs over 10x, and model architecture is the only math lever, argues thread — ChengleiSi · 2026-09-10
- DeepSeek unveils asymmetric Causal Encoder-Decoder: 552B MoE with just 8B active input params — ChengleiSi · 2026-09-10
- Switch Transformer by hand: a 13-step walkthrough of how sparse MoE works — ProfTomYeh · 2026-09-10