Local test: Qwen 3.8 27B's overthinking brings it near Sonnet-level performance
maxwell321 · reddit · 2026-08-17
A user running local models on 3x 3090s plus a Tesla P40 reviewed Qwen 3.8 27B (Unsloth UD-Q8KXL quant), concluding its heavy thinking draws out fine details from training knowledge, landing near Claude Sonnet with potential for Opus-level results.
The benchmark: 1:1 recreations of classic arcade games. Qwen 3.6's Galaga clone was effectively a Space Invaders clone — no diving enemies, no shooting, missing details requiring hand-holding. Qwen 3.8 built sprites from dynamic pixel bitmaps with two-frame animation, added a CRT filter and power-on simulation, and even remembered the fighter-capture mechanic (rescue for dual ships), only replacing the capture beam with collision. The author says this closes much of the local-vs-proprietary gap, worth the extra inference time and context.
More from Models
- ChatGPT's 1M-token context can be unlocked, but requests over 272K tokens cost more — SimplyAnnisa · 2026-08-17
- Dev on OpenAI 1M context: Seamless compaction is the real win — eyishazyer · 2026-08-17
- Users report Qwen models "overthink" simple responses, exhausting context — malliktwts · 2026-08-17
- Local benchmark: Qwen 3.8 and DeepSeek 4 show significant quality jump on Mac — olcan · 2026-08-17
- Sarvam AI Unveils 105B Voice Model and Full-Stack Agent Solutions — AashaySachdeva · 2026-08-17
- GPT-4o Pro run time exceeds 2 hours in user test — banteg · 2026-08-17