105 planted bugs tested: local Qwen3.8-27B nearly matches Claude Opus at bug fixing

PawelHuryn · x · 2026-09-17

The author benchmarked local LLMs on real work: two repos with 105 planted bugs to find and fix. Results: Opus 4.8 (max) fixed 15, Qwen3.8-27B 8-bit fixed 10.7, Sonnet 5 (high) 9, Gemma 4 31B only 4, and gpt-oss-120b scored 0 (useless).

The surprise was the 8-bit Qwen3.8-27B: at 28.6 GB, it runs comfortably on a 64 GB Mac mini with a 200K context window. Not as smart as frontier closed models, but capable of real, useful bug-fixing work at the cost of electricity alone.

Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→

Original post →

More from Models

Models channel →