Qwen 3.8 27B thinks until context runs out, never produces output locally
WyattTheSkid · reddit · 2026-08-20
A user reports that running Qwen 3.8 27B (bf16 then Q8) locally via LM Studio/llama.cpp with opencode, with thinking set to medium and official sampling params, the model just keeps reasoning until it exhausts the context and never produces a deliverable — and setting an 8192 reasoning budget makes it not respond at all. Asking the community for settings and workarounds.
Related event: Qwen 3.8 27B Overflows Terminal with Endless Thinking Locally(2 posts)→
More from Models
- llama.cpp dflash2: Qwen 3.8 27B Inference Speed Up to 3x — Top-Eye-8104 · 2026-08-20
- DeepSeek V4 Pro Challenges Closed-Source Models with GPT-4o Level Performance at Lower Cost — Two Minute Papers · 2026-08-20
- Dev reports: Qwen 3.8 27B only works with reasoning_level set to low — andrejusb · 2026-08-20
- Grok's free tier is real but unstable with changing caps — heypearlai · 2026-08-20
- Hack: Get free uncapped access to 400+ models via LMArena — heypearlai · 2026-08-20
- ChatGPT free goes unlimited; Claude free tier offers Sonnet 5 — heypearlai · 2026-08-20