Qwen 3.8 27B local coding session runs 3 days on one RTX 4090, then spews endless slashes

Tiny-Entertainer-346 · reddit · 2026-09-23

The author ran Qwen3.8-27B (UD Q4KXL GGUF) via llama-swap on a single RTX 4090 with 32GB RAM inside VS Code Copilot, keeping one chat session alive for 3 days — over 20k lines of text, 131072 context, Q8 KV cache. Prefill ran at 2200-2800 tok/s and generation 23 t/s. Now llama-server logs look normal but the chat response is an endless string of slashes; they're asking why.

Original post →

More from Infra

Infra channel →