Bug Hunt Bench grades frontier coding models on 105 real bugs in production repos

PawelHuryn · x · 2026-09-13

Developer Pawel Huryn released Bug Hunt Bench: 105 real bugs planted in two production codebases, with frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, GLM and more) getting one round per repo in their own agentic CLIs (Codex CLI, Claude Code, Grok CLI, Antigravity CLI). Every diff is blind-graded against a withheld answer key, counting only planted bugs. Live leaderboard includes method notes and every receipt. Early finding: Muse Spark 1.3's xhigh mode is 50% pricier and 13% slower than high but fixes just one more bug — max effort seems to buy the biggest single jump.

Related event: Bug Hunt Bench: Muse Spark 1.3 Ties for First, Meta Nears Frontier(4 posts)→

Original post →

More from coding & agent

coding & agent channel →