Bug Hunt Bench: frontier coding models graded on 105 planted real-world bugs, cost spread ~200x
PawelHuryn · x · 2026-09-24
Pawel Huryn (The Product Compass) launched Bug Hunt Bench, a hands-on leaderboard grading frontier coding models on 105 bugs planted in real repos.
- Blind grading: one prompt per repo, blind-graded scoring; supports comparing the same model across reasoning tiers and harnesses.
- Cost view: includes score-vs-cost and score-vs-time charts — cost spread is roughly 200x (log scale), while runtime spans under 7x (linear scale).
- Full run notes and caveats are public on GitHub (results/run-notes.md); the author promises hands-on, no-hype, run-first methodology.
More from coding & agent
- They audited 13 Reddit MCP servers: 97 tools, none handle what happens after posting — investigatormaker · 2026-09-24
- AI agent acts as F1 TV director with LangChain, Nemotron, and LangSmith — thetripathi58 · 2026-09-24
- huggingface_hub v1.33.0 ships: hf skills add covers Claude Code, Codex, Cursor — vanstriendaniel · 2026-09-24
- Where an agent lives matters more than what model it runs, says 6-month test — thirdtea4 · 2026-09-24
- AI agent works fine 95% of the time — the problem is the other 5% — Darede_ · 2026-09-24
- Stealth Model Space Bunny Free on OpenCode: One-Prompt Full Game, 1M Context, Zero Retention — iamfakhrealam · 2026-09-24