Bug Hunt Bench leaderboard live: 105 planted bugs, score vs cost across frontier models
PawelHuryn · x · 2026-09-24
Pawel Huryn published the full Bug Hunt Bench leaderboard: blind-graded runs against 105 planted bugs in real repos, one prompt per repo, with score-vs-cost (log axis, roughly 200x spread) and score-vs-time (linear, under 7x) views. Full run notes and caveats live on GitHub in results/run-notes.md. The board is maintained alongside his newsletter The Product Compass.
Related event: Bug Hunt Bench Leaderboard Released; Code-Review Skill Boosts Model Scores(2 posts)→
More from coding & agent
- Solo Dev Builds MCP Server So Claude Code Ships APKs Straight to Phones — justabigmilkShake · 2026-09-24
- Seroter Daily #873: The Internet Isn't Ready for the Agentic Wave, MCP Apps, GKE Scale-to-Zero — rseroter · 2026-09-24
- Krea Agent adds custom apps: build creative tools by describing them — angrypenguinPNG · 2026-09-24
- Hitting $200 Claude and Codex rate limits: this GPT-6 + GLM 5.3 Flash combo costs 4x less — Cole Medin · 2026-09-24
- qontoctl: A MCP Server for Qonto Banking Accounts Lands on Glama — modelcontextprotocol · 2026-09-24
- Gemini CLI v0.61.0 ships indirect prompt injection fix and sandbox hardening — gemini-cli-robot · 2026-09-24