Qwen3.8-Flash-Next Stuck in Endless Thinking at Low Reasoning Level

pabloodiablo · reddit · 2026-08-30

A developer reported severe performance issues with Qwen3.8-Flash-Next (Q8 quantized) on a 2x 128GB VRAM cluster. Even with the reasoning effort parameter set to low, the model entered an "endless thinking" loop during code generation, running for over 45 minutes without stopping (reaching 32k context), while generating significantly more output than comparable models (Qwen3.8-27B and DeepSeek v4 both finished in under 15 minutes). The author suspects a llama.cpp configuration issue or a bug in the model and is seeking community input.

Original post →

More from Models

Models channel →