Qwen3.8-Flash-Next Stuck in Endless Thinking at Low Reasoning Level
pabloodiablo · reddit · 2026-08-30
A developer reported severe performance issues with Qwen3.8-Flash-Next (Q8 quantized) on a 2x 128GB VRAM cluster. Even with the reasoning effort parameter set to low, the model entered an "endless thinking" loop during code generation, running for over 45 minutes without stopping (reaching 32k context), while generating significantly more output than comparable models (Qwen3.8-27B and DeepSeek v4 both finished in under 15 minutes). The author suspects a llama.cpp configuration issue or a bug in the model and is seeking community input.
More from Models
- Celeris-1 Magnus: New Model Claims Top Spot on τ³-bench for Agentic Work — timshi_ai · 2026-09-01
- Focus on specific tasks, not the best model, as selection logic evolves — aftahi_ai · 2026-09-01
- User Rants on GPT-5.6 Hallucinations and Coding Limits, Hopes for GPT-6 Fix — Prestigiouspite · 2026-09-01
- Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization — alejandroll10 · 2026-09-01
- Rumor: GPT-6 'Astra' nears human-level computer use — jYtanYj · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01