Users report Qwen3.8 outputting garbage after prolonged use
trashacct383 · reddit · 2026-08-27
A user reports that Qwen3.8-27B-FP8 starts outputting nonsense or garbage after a few hours of use on vLLM, requiring a restart to fix. The user provided detailed configuration flags (including FlashInfer backend, xgrammar, MTP speculative decoding) and noted that adding a repetition penalty helped slightly, but vLLM 28 seemed to worsen the issue, asking if others are experiencing the same.
More from Infra
- Nvidia projects 70% revenue growth for fiscal 2028, beating analyst expectations of 44% — firstadopter · 2026-08-27
- Nvidia's NVHBM Brings 30% More Bandwidth to NVLink Fusion; Amazon Annapurna First Partner — nordicinst · 2026-08-27
- AWS and NVIDIA to Deploy 2 Million Additional GPUs for Agentic and Physical AI — nvidia · 2026-08-27
- Gemini 3.5 Transcribe Now Available on Vercel AI Gateway — osanseviero · 2026-08-27
- Nvidia guides $108B Q3 revenue, doubling growth even without China data center sales — inductionheads · 2026-08-27
- Palantir Karp: Serious enterprises need to own their models and infrastructure — JosephJacks_ · 2026-08-27