Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points
PawelHuryn · x · 2026-09-15
Pawel Huryn ran hands-on blind-graded evaluations on Bug Hunt Bench (105 real planted bugs, one prompt per repo) using his Muse Code subscription, with costs expressed as API equivalents.
Key findings:
- Multiple runs significantly increase bug detection, especially for smaller models at lower reasoning effort
- For stronger models at higher effort levels, extra runs barely move the average (max 1-2 points)
- Cost across runs spans roughly 200x (plotted on a log axis); time spans under 7x
- The leaderboard supports sorting by score, cost, time, and coverage, with full run notes on GitHub
Takeaway: small models can approach frontier single-run performance via multiple passes, but paying for a stronger config beats re-running frontier models.
Related event: Bug Hunt Bench Testing: Multi-Run Boosts Bug Detection in Smaller Models(2 posts)→
More from coding & agent
- AI agent swarm beats NanoChat benchmark SoTA in 3 days with 15k-node knowledge graph — hyperparticle · 2026-09-15
- A/B testing the i-have-adhd plugin: making coding agent answers scannable — KhuyenTran16 · 2026-09-15
- rekursiv.ai open-sources trackinizer, an epistemological database for agent research — hyperparticle · 2026-09-15
- Give your coding agent GPUs via the Hugging Face CLI in one link — ben_burtenshaw · 2026-09-15
- Hybrid search put to the test: dense embeddings, BM25 and SPLADE on 47 SEC filings in Qdrant — qdrant_engine · 2026-09-15
- Agent running on Opus nearly clones an ElevenLabs ad video — can you spot the difference? — austin_malerba · 2026-09-15