Bug Hunt Bench: Qwen3.8-27B run 3x beats single Opus 4.8 max on planted bug detection

PawelHuryn · x · 2026-09-17

Pawel Huryn launched Bug Hunt Bench, blind-testing frontier coding models on 105 real repos with planted bugs. Key findings:

The board is blind-graded with one prompt per repo; data and caveats are public on GitHub, and only runs the author executed are listed.

Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→

Original post →

More from Models

Models channel →