Qwen's Overthinking is a Feature, Adjust Your Expectations

wombweed · reddit · 2026-08-18

Addressing complaints about Qwen 3.8-27B's long thinking times, the author argues that if speed is critical, users should disable thinking or reduce the reasoning effort. If loops occur, higher quantization may help. The author emphasizes that extensive thinking is a result of RL training for verification and deliberation, which contributes to the model's high benchmark scores. At the 27B scale, users shouldn't expect instant encyclopedic knowledge like massive models; instead, they should leverage the model's solid fundamentals and consistent problem-solving approach.

Related event: Qwen 3.8 27B "Overthinking" Debate: Tests Point to Necessary Cost for Performance(7 posts)→

Original post →

More from Models

Models channel →