Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points

PawelHuryn · x · 2026-09-15

Pawel Huryn ran hands-on blind-graded evaluations on Bug Hunt Bench (105 real planted bugs, one prompt per repo) using his Muse Code subscription, with costs expressed as API equivalents.

Key findings:

Takeaway: small models can approach frontier single-run performance via multiple passes, but paying for a stronger config beats re-running frontier models.

Related event: Bug Hunt Bench Testing: Multi-Run Boosts Bug Detection in Smaller Models(2 posts)→

Original post →

More from coding & agent

coding & agent channel →