Bug Hunt Bench: 105 real planted bugs, GPT-6 Astra leads at 45, Pareto hits 30.7 for just $4.81
PawelHuryn · x · 2026-09-18
Pawel Huryn's Bug Hunt Bench pits frontier coding models against 105 real bugs planted in 2 repos (all bugs frontier models initially failed on), blind-graded with no same-lab peer scoring.
- Score vs cost: GPT-6 Astra (max): 45, $33.03; Fable 5.1 (max): 43, $77.55; Muse Spark 1.3 (max): 32.2, $18.11; Pareto (Unbiased, ex-Union Alpha): 30.7, $4.81; Grok 4.6 (xhigh): 28.7, $18.60; Opus 5 (max): 27, $51.33
- Pareto stands out on value: 30 points for under $5 — and turns out to be a composite model
- Cost spread across the board is 200x (log axis); time spread under 7x
- Live leaderboard, data and run notes are public
Takeaway: cheap newer models already approach the most expensive frontier models on real bug-fixing, though the priciest still set the ceiling.
Related event: Pareto Matches Top Models on 105-Bug Benchmark for $4.81(2 posts)→
More from coding & agent
- ChatGPT web can now open GitHub PRs directly, despite outdated docs saying it can't — dotey · 2026-09-18
- Sakana AI Launches Fugu Max, a Multi-Agent Orchestrator Routing Tasks to Leanest Capable Models — tkasasagi · 2026-09-18
- No Java SDK needed: one plain Java 25 file to call TypeSafe's Jev API — therealdanvega · 2026-09-18
- fast-jev-compaction: Claude Code plugin cuts a ~1M-token session to 86K in 1 second — viksit · 2026-09-18
- Claude Code 2.1.276 by the numbers: shipped in under 6 hours, +440 prompt tokens — ClaudeCodeLog · 2026-09-18
- Claude Code 2.1.276 fixes regression breaking all requests behind proxies — ClaudeCodeLog · 2026-09-18