Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing
A new benchmark with 105 planted bugs in two real codebases shows that locally-run Qwen3.8-27B fixes nearly as many bugs as Claude Opus 4.8, and beats a single Opus run across three attempts.
2026-09-17 ~ 2026-09-17 · 3 related posts
- 105 planted bugs tested: local Qwen3.8-27B nearly matches Claude Opus at bug fixing — PawelHuryn · 2026-09-17
- Planted 105 bugs in 2 repos: Qwen3.8-27B 8-bit run 3x beats a single Opus 4.8 run — PawelHuryn · 2026-09-17
- Bug Hunt Bench: Qwen3.8-27B run 3x beats single Opus 4.8 max on planted bug detection — PawelHuryn · 2026-09-17