Qwen 3.8 27B local coding session runs 3 days on one RTX 4090, then spews endless slashes
Tiny-Entertainer-346 · reddit · 2026-09-23
The author ran Qwen3.8-27B (UD Q4KXL GGUF) via llama-swap on a single RTX 4090 with 32GB RAM inside VS Code Copilot, keeping one chat session alive for 3 days — over 20k lines of text, 131072 context, Q8 KV cache. Prefill ran at 2200-2800 tok/s and generation 23 t/s. Now llama-server logs look normal but the chat response is an endless string of slashes; they're asking why.
More from Infra
- Cloudflare CTO Dane Knecht makes TIME's 2026 executives list as AI crawlers hit 52% of traffic — dinasaur_404 · 2026-09-23
- Rat Stack: Build Your App and Cloud as One Typed Program So Agents Deploy Reliably — samgoodwin89 · 2026-09-23
- optimAIzr: A local-first CLI that audits AI token waste, works with Claude Code and Codex — stichstichstich · 2026-09-23
- 256GB Mac Studio reselling for $6,000 over MSRP amid local AI demand — GabGarrett · 2026-09-23
- Inside the AI Factory: How CPUs and Memory Limit GPU Inference at Scale — BenBajarin · 2026-09-23
- Epoch AI: The plunging price of thought — intelligence keeps getting radically cheaper — Proper_Actuary2907 · 2026-09-23