Bug Hunt Bench: Qwen3.8-27B run 3x beats single Opus 4.8 max on planted bug detection
PawelHuryn · x · 2026-09-17
Pawel Huryn launched Bug Hunt Bench, blind-testing frontier coding models on 105 real repos with planted bugs. Key findings:
- The bugs are tough ones frontier models missed at the start of 2026, not random trivia — yet Qwen reviewing a normal PR catches far more
- Running Qwen3.8-27B (8-bit) three times beats a single run of Opus 4.8 (max) in detection coverage, showing the payoff of repeated sampling for bug hunting
- Costs across the leaderboard span roughly 200x (log scale); runtime under 7x
The board is blind-graded with one prompt per repo; data and caveats are public on GitHub, and only runs the author executed are listed.
Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→
More from Models
- "Direct confidence readout" claim debunked: it's just entropy from the logit distribution — mgostIH · 2026-09-17
- Claude flags Kali Linux install as cyber, sparking over-refusal backlash — TinfoilTricorn · 2026-09-17
- Fixing one simple grader doubled the score—and the rest were errors in the tasks — xeophon · 2026-09-17
- Claude flagged a routine Kali Linux install as cyber, sparking overrefusal complaints — TinfoilTricorn · 2026-09-17
- Rumor: Gemini 4.0 Pro spotted testing in LMArena under the name Gemini 3.8 Flash — gaganghotra_ · 2026-09-17
- $200 plans are just a preview: analyst predicts frontier models will go API-only — StewartalsopIII · 2026-09-17