Qwen3.5 Still Performs Extended Reasoning Even When Thinking Mode is Disabled
_lewtun · x · 2026-07-30
Hugging Face engineer Lewis Tunnicliffe noticed that when evaluating Qwen3.5 models with enablethinking=false, the models still exhibit a form of extended reasoning, generating "Wait, ..." or even <think> tags. This is particularly evident in out-of-domain evaluations (e.g., biology), resulting in massive token generation. In contrast, Gemma4 models strictly follow the setting and produce concise outputs, raising questions about their differing post-training recipes.
More from Models
- LightOn Releases Multilingual Retrieval Models and 2.8B Training Pairs — IgorCarron · 2026-07-30
- P-Image-Ideogram Hits Pareto Frontier for Image Gen Speed and Cost — _akhaliq · 2026-07-30
- Even Gemini is Smarter Than a Middle Schooler — DangerousSpray3656 · 2026-07-30
- MoTA: Replace Massive Context with 4MB LoRA Adapters, Cutting Inference Storage by 100x — EyalToledano · 2026-07-30
- LightOn Releases Fully Open Multilingual Retrieval Models mDenseOn & mLateOn — antoine_chaffin · 2026-07-30
- Kimi K3 is Good but Pricey: First to Cut Costs Wins $10B Valuation — bindureddy · 2026-07-30