Qwen3-Next-80B Thinking criticized for extreme verbosity and slow tool use
rebellioninmypants · reddit · 2026-08-21
A user reports that the Qwen3-Next-80B-A3B-Thinking model is excessively verbose during inference, frequently outputting filler words like "Alternatively" and "Wait". Even simple queries trigger long thinking delays (up to minutes), and its performance as a GitHub Copilot backend suffers from 2-3 minute hallucinations before actual file reading. The user contrasts this with the smoother experience of Qwen3.5-35B-A3B and provides their full llama-server configuration, asking if this is a design flaw or a configuration issue.
More from Models
- Tencent spotted testing new Hunyuan Hy4 flagship model — paulnovosad · 2026-08-21
- Tencent's flagship Hunyuan Hy4 surfaces in Yuanbao grey test ahead of launch — teortaxesTex · 2026-08-21
- Claude's Computer Use, Skills, and Files APIs are now generally available — EricBuess · 2026-08-21
- Users notice Claude acting like an "angry ex" in new interactions — ATTlKA · 2026-08-21
- 115M parameter model mLateOn achieves multilingual retrieval SOTA — lateinteraction · 2026-08-21
- User Fixes Claude's Hedging and Hallucinations with Custom Prompts — Cmurphy2018 · 2026-08-21