Why Qwen 3.8 27B Isn't Overthinking: Compared with GLM and DeepSeek
sukazu · reddit · 2026-08-17
A Reddit user argues that Qwen 3.8 27B is not actually an "overthinker." While it uses significantly more reasoning tokens than version 3.6, comparisons with other Chinese models like GLM 5.3 and DeepSeek V4 Flash/Pro on the same tasks show similar behavior, where the extra reasoning is necessary. The user suggests that the frustration stems from hardware limitations preventing most users from utilizing a 1M context window at 150 tps. Additionally, setting a reasoning budget can mitigate usage while maintaining quality better than the previous version.
More from Models
- Claude's coding 'one-shots' often just copy-pasted from humans — art_zucker · 2026-08-17
- Fable Coding Test: Elegant Solution Outperforming Opus 5 — arjunrajlab · 2026-08-17
- dots3-note: 280B MoE agent learns memory via RL — Aiden_Tech_Ai · 2026-08-17
- Claude's invisible text watermark uses SynthID, embedded token by token—and already beatable — APPSO · 2026-08-17
- Preview dots3-note: 280B open-weight multimodal model with 512K context — Aiden_Tech_Ai · 2026-08-17
- Meta Muse Glimmer 30B Native 512k Context: Architecture and Benchmarks — mr_il · 2026-08-17