Bug Hunt Bench Testing: Multi-Run Boosts Bug Detection in Smaller Models

Pawel Hur ran blind tests of frontier coding models on Bug Hunt Bench (105 real injected bugs), finding multi-run significantly improves bug detection for smaller models while strong models gain only 1-2 points; he later clarified Muse Spark 1.3 pricing uses standard API rates.

2026-09-15 ~ 2026-09-15 · 2 related posts