Qwen 3.8 Next Flash is painfully verbose: 13-minute thinking on single coding prompts
Infinite-Local5435 · reddit · 2026-09-07
A long-time Qwen 3.6 27B user on Reddit reports that Qwen 3.8 Next Flash is extremely verbose: with 150 tokens/sec generation, a single coding request can spend 13 minutes thinking (7000 tokens).
- Output is decent most of the time, but any task requiring decision-making degenerates into buzzword soup nobody in SWE actually uses
- BYOK into VSCode runs faster, but via pi.dev a request can take up to an hour
- The author avoids lowering the thinking level since no solid benchmarks exist for performance at different levels, and 3.8 27B suggests quality drops noticeably
More from Models
- Astra reportedly much better at musical reasoning and geometry, anecdotally — latticecut · 2026-09-07
- Mystery model Omen Alpha spotted; tokenizer tests point to new Zhipu GLM — realsohamparekh · 2026-09-07
- Training mixtures are now all synthetic: small-model training is really distillation — RexDouglass · 2026-09-07
- Sol High usage test: one complex prompt eats 5% of the 5-hour limit — remixedmoon5 · 2026-09-07
- Bodhan AI open-weights speech, vision and translation models for Indian languages on Hugging Face — selfawareatom · 2026-09-07
- VoiceMem: A Dual-Brain Memory System Cuts Voice AI Retrieval to 134ms — 机器之心 · 2026-09-07