Bug Hunt Bench ranks GPT-6 Astra top as coding models fix real planted bugs, costs spread 200x
PawelHuryn · x · 2026-09-23
Pawel Huryn launched Bug Hunt Bench, a blind-graded leaderboard pitting frontier coding models against 105 real planted bugs (one prompt per repo, judges from different model families scoring against a secret answer key).
- Current ranking: GPT-6 Astra > GPT-5.6 Sol ≈ Opus 5.5 > Fable 5.1 > Opus 5
- Only pre-planted bugs count; unplanted bug reports (notably from OpenAI's "bugmaxing") are excluded, which the author argues is arguably negative
- Costs across runs span roughly 200x (log axis); time spread is under 7x
- Next up: n=3 run for Opus 5.5 and a separate GPT-6 Sol post
Related event: Bug Hunt Bench: GPT-6 Astra Leads, MiMo Shines on Cost, Grok 4.7 Stalls(14 posts)→
More from coding & agent
- ChatGPT co-inventor's notes: a 100ms context filter in front of Claude Code to kill compaction — PrajwalTomar_ · 2026-09-23
- Alchemy adds Kubernetes docs hub: kind-to-EKS tutorial covering Deployments, Helm and server-side apply — samgoodwin89 · 2026-09-23
- Rat Stack: Build Your App and Cloud as One Typed Program So Agents Deploy Reliably — samgoodwin89 · 2026-09-23
- 'Remember tab complete?' How fast AI coding paradigms have moved — yoobinray · 2026-09-23
- Survey: One in Four AI Agents Goes Unmonitored Despite 12 Observability Tools — rseroter · 2026-09-23
- Massive x402 Lets AI Agents Pay for Live Market Data via Coinbase — kleffew94 · 2026-09-23