Qwen 3.8 27B thinks until context runs out, never produces output locally

WyattTheSkid · reddit · 2026-08-20

A user reports that running Qwen 3.8 27B (bf16 then Q8) locally via LM Studio/llama.cpp with opencode, with thinking set to medium and official sampling params, the model just keeps reasoning until it exhausts the context and never produces a deliverable — and setting an 8192 reasoning budget makes it not respond at all. Asking the community for settings and workarounds.

Related event: Qwen 3.8 27B Overflows Terminal with Endless Thinking Locally(2 posts)→

Original post →

More from Models

Models channel →