Qwen 3.8 27B fails complex coding task, lags behind DeepSeek and GLM
myreala · reddit · 2026-08-19
A user reported a disappointing experience with Qwen 3.8 27B on a complex C kernel coding task. The objective was to adjust thread count limits in a TTS execution kernel originally written by DeepSeek Pro. Despite running on a single 3090 with 150k context and fp8 kv cache for about six hours across multiple sessions, Qwen entered a loop, produced unusable code, and suffered from tool call failures. In contrast, GLM 5.3 completed the entire task in about twenty minutes. The user expressed reluctance to invest in higher quantization or hardware given the poor results.
More from Models
- Users Complain Claude IQ Drop, Anthropic Scrambles to Recover — annetgriffin · 2026-08-19
- User claims Flash 3.7 beats GPT-Terra and Claude as Google's best model — bindureddy · 2026-08-19
- User Review: Gemini 3.7 Flash rated as best model for internal Slack bot — amankhan · 2026-08-19
- GLM 5.3 ranks among leaders in new benchmarks — markjeffrey · 2026-08-19
- Dev shows Qwen3.8 its own README and the model guesses why it performs so well — MikePFrank · 2026-08-19
- Chart reveals Qwen 3.8 as a massive performance outlier for its size — MikePFrank · 2026-08-19