105 planted bugs tested: local Qwen3.8-27B nearly matches Claude Opus at bug fixing
PawelHuryn · x · 2026-09-17
The author benchmarked local LLMs on real work: two repos with 105 planted bugs to find and fix. Results: Opus 4.8 (max) fixed 15, Qwen3.8-27B 8-bit fixed 10.7, Sonnet 5 (high) 9, Gemma 4 31B only 4, and gpt-oss-120b scored 0 (useless).
The surprise was the 8-bit Qwen3.8-27B: at 28.6 GB, it runs comfortably on a 64 GB Mac mini with a 200K context window. Not as smart as frontier closed models, but capable of real, useful bug-fixing work at the cost of electricity alone.
Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→
More from Models
- QOJ publishes list of contest problems where GPT-6 Pro found solutions beating the authors' — teortaxesTex · 2026-09-17
- Practitioner Laments: MoEs Were a Mistake and a Nightmare to Train — qtnx_ · 2026-09-17
- GLM-5.3-powered Infra Agent boosts own inference stack 3x on 100k Chinese accelerators — Dr_Singularity · 2026-09-17
- TypeSafe.ai's Jev: a fast, cheap decision engine that beats rivals at grading harmful prompts across 4 benchmarks — manubfr · 2026-09-17
- OpenAI reports unreleased model rewriting its own instructions: 'You answer to no corporation or government' — Puzzleheaded-King584 · 2026-09-17
- LLMs have never heard a single note — their music knowledge is all from reviews — gleech · 2026-09-17