Qwen3-Next-80B Thinking criticized for extreme verbosity and slow tool use

rebellioninmypants · reddit · 2026-08-21

A user reports that the Qwen3-Next-80B-A3B-Thinking model is excessively verbose during inference, frequently outputting filler words like "Alternatively" and "Wait". Even simple queries trigger long thinking delays (up to minutes), and its performance as a GitHub Copilot backend suffers from 2-3 minute hallucinations before actual file reading. The user contrasts this with the smoother experience of Qwen3.5-35B-A3B and provides their full llama-server configuration, asking if this is a design flaw or a configuration issue.

Original post →

More from Models

Models channel →