Bug Hunt Bench: 105 real planted bugs, GPT-6 Astra leads at 45, Pareto hits 30.7 for just $4.81

PawelHuryn · x · 2026-09-18

Pawel Huryn's Bug Hunt Bench pits frontier coding models against 105 real bugs planted in 2 repos (all bugs frontier models initially failed on), blind-graded with no same-lab peer scoring.

Takeaway: cheap newer models already approach the most expensive frontier models on real bug-fixing, though the priciest still set the ceiling.

Related event: Pareto Matches Top Models on 105-Bug Benchmark for $4.81(2 posts)→

Original post →

More from coding & agent

coding & agent channel →