Bug Hunt Bench: 105 real bugs stress-test GPT-6, Claude, Grok, Gemini and more coding agents
PawelHuryn · x · 2026-09-05
Developer Pawel Huryn launched Bug Hunt Bench (GitHub: phuryn/bug-hunt-bench) with a live leaderboard, website and full data.
Design highlights:
- 105 real bugs hidden across two production codebases
- Tested frontier coding models: GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, GLM and more
- Each model gets one round per repo in its own agentic CLI (Codex CLI, Claude Code, Grok CLI, Antigravity CLI) to find and fix bugs
- Every diff is graded blind against a withheld answer key; only planted bugs count, with every receipt published
- Additional effort-level results and newly released models will be added over time
The goal: save others the cost of running their own benchmark.
More from coding & agent
- xAI to host Grok Bot Galaxy, a three-day agent event in SF this September 15-17 — XFreeze · 2026-09-05
- Developer who found Codex 'slow and annoying' 1.5 years ago declares AGI achieved — gajesh · 2026-09-05
- OpenAI forum report paints agents roaming the internet like raiding nomad hordes — tedmitew · 2026-09-05
- One prompt builds a playable Naruto 2D action game to test ChatGPT's limits — callme_e · 2026-09-05
- skill-builder: open-source tool turns workflow descriptions into production-ready Codex skills — cneuralnetwork · 2026-09-05
- AI-generated Minecraft-like game adds floods, tornadoes and a Death Star, open source soon — ChrisGPT · 2026-09-05