Bug Hunt Bench Testing: Multi-Run Boosts Bug Detection in Smaller Models
Pawel Hur ran blind tests of frontier coding models on Bug Hunt Bench (105 real injected bugs), finding multi-run significantly improves bug detection for smaller models while strong models gain only 1-2 points; he later clarified Muse Spark 1.3 pricing uses standard API rates.
2026-09-15 ~ 2026-09-15 · 2 related posts
- Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points — PawelHuryn · 2026-09-15
- Bug Hunt Bench author clarifies Muse Spark 1.3 costs use standard API pricing — PawelHuryn · 2026-09-15