vLLM deployment of Qwen 3.8 27B suffers from endless reasoning loops

germangrower69 · reddit · 2026-08-16

A developer reports severe latency issues deploying Qwen 3.8 27B via vLLM on an RTX 6000 Pro. Despite setting 'thinking effort' to low, medium, or xhigh, the model enters excessively long reasoning phases (1-5 mins), whereas Qwen 3.6 and DeepSeek V4 Flash respond in 20-30 seconds on the same hardware. Extensive tweaking of quantization, flags, and recipes has failed to resolve the issue.

Original post →

More from Infra

Infra channel →