Planted 105 bugs in 2 repos: Qwen3.8-27B 8-bit run 3x beats a single Opus 4.8 run
PawelHuryn · x · 2026-09-17
PawelHuryn tested local LLMs on real work: 2 repos with 105 planted bugs, asking models to find and fix what they can. Single-run scores: Opus 4.8 (max) found 15, Qwen3.8-27B 8-bit 10.7, Sonnet 5 (high) 9, Gemma 4 31B only 4, and gpt-oss-120b 0 (useless).
The surprise was Qwen3.8-27B: the 8-bit model is about 28.6 GB and comfortably runs on a 64 GB Mac mini with a 200K context window. Running it three times beats a single maxed-out Opus 4.8 run — a strong data point for local models on real engineering tasks.
Related event: Bug Hunt Bench: local Qwen3.8-27B nears Claude Opus in bug fixing(3 posts)→
More from coding & agent
- QOJ publishes list of contest problems where GPT-6 Pro found solutions beating the authors' — teortaxesTex · 2026-09-17
- What Happens After You Tell an AI Agent It's Wrong? Dreamforce Enterprise Lessons — TheTuringPost · 2026-09-17
- GitHub MCP maintainer packs the house at MCPCon with talk "MCP doesn't have a context problem" — marlene_zw · 2026-09-17
- Label the Row: A Six-Step Data Classification Cheat Sheet for AI Products — blaizedsouza · 2026-09-17
- Agent Swarms as the Third Scaling Axis: 10,000-Agent Test Run Proves It Works — Upset-Winter7174 · 2026-09-17
- What happens when a RAG agent retrieves a poisoned document? A reusable security test case — Tophant_ · 2026-09-17