Bug Hunt Bench: Benchmarking frontier coding models on real planted bugs

PawelHuryn · x · 2026-08-26

Pawel Huryn released Bug Hunt Bench, a live benchmark for frontier coding models. It tests bug-fixing abilities by planting real bugs in repositories, with a leaderboard of 105 bugs. The platform ranks models by bugs fixed, cost (USD), and time (minutes), featuring blind grading and single-prompt constraints to evaluate real-world debugging cost-effectiveness.

Related event: Bug Hunt Bench Launches to Test Real Bug Fixing by Code Models(2 posts)→

Original post →

More from coding & agent

coding & agent channel →