Planted 105 bugs in 2 repos: Qwen3.8-27B 8-bit run 3x beats a single Opus 4.8 run

PawelHuryn · x · 2026-09-17

PawelHuryn tested local LLMs on real work: 2 repos with 105 planted bugs, asking models to find and fix what they can. Single-run scores: Opus 4.8 (max) found 15, Qwen3.8-27B 8-bit 10.7, Sonnet 5 (high) 9, Gemma 4 31B only 4, and gpt-oss-120b 0 (useless).

The surprise was Qwen3.8-27B: the 8-bit model is about 28.6 GB and comfortably runs on a 64 GB Mac mini with a 200K context window. Running it three times beats a single maxed-out Opus 4.8 run — a strong data point for local models on real engineering tasks.

Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→

Original post →

More from coding & agent

coding & agent channel →