Bug Hunt Bench grades frontier coding models on 105 real bugs in production repos
PawelHuryn · x · 2026-09-13
Developer Pawel Huryn released Bug Hunt Bench: 105 real bugs planted in two production codebases, with frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, GLM and more) getting one round per repo in their own agentic CLIs (Codex CLI, Claude Code, Grok CLI, Antigravity CLI). Every diff is blind-graded against a withheld answer key, counting only planted bugs. Live leaderboard includes method notes and every receipt. Early finding: Muse Spark 1.3's xhigh mode is 50% pricier and 13% slower than high but fixes just one more bug — max effort seems to buy the biggest single jump.
Related event: Bug Hunt Bench: Muse Spark 1.3 Ties for First, Meta Nears Frontier(4 posts)→
More from coding & agent
- Codex denies hash pinning blame, admits 10 minutes later it was exactly that — moyix · 2026-09-13
- 'You don't need math to be an AI engineer': a 6-step hiring playbook goes viral — ashishllm · 2026-09-13
- Codex script bulk-downloads 9,000 baby monitor photos in 45 minutes — thegautamkamath · 2026-09-13
- 'The End of Prompt Engineering': 300 Kimi K3 agents, one AGENTS.md, zero escapes in 41 days — JohnAlexander · 2026-09-13
- Reddit user asks: one Astra prompt eats the whole weekly quota—what's wrong with my workflow? — Cetnet · 2026-09-13
- Automation's blind spot: the undocumented tasks humans quietly do — chris_j_paxton · 2026-09-13