Running Qwen3.8 Flash Next on 128GB RAM + one 5080: 136pp/19tg at Q5_K_XL
whatyathinkk · reddit · 2026-09-15
A detailed local-inference setup for Qwen3.8 Flash Next on 128GB RAM + a single RTX 5080 (16GB): Unsloth's Q5KXL hits 136pp/19tg, Q4KXL 152pp/23tg, Q3KXL 195pp/23tg. Slow but impressive quality, the poster says.
Config highlights: 6-part GGUF (UD-Q5KXL-ncmoe48), lazy-mode on, ngl=999 with n-cpu-moe=48 keeping MoE layers on CPU, perlayertokenembd forced to CPU, 262K context, q80 KV cache, flash-attn on; sampling at temp=1.0/top-k=20/top-p=0.95 with thinking preserved and reasoningeffort=medium. Full llama.cpp config included; the poster asks about MTP or other speedups.
More from Infra
- OpenAI engineers: kernel optimization cut GPT-5.6 Sol serving cost by 20% — TheTuringPost · 2026-09-15
- Devin left alone with Modal H100s cuts training kernel peak memory 46% and latency 52% — AAAzzam · 2026-09-15
- Why Macs quietly win at local AI: unified memory beats RTX 5090 and accessibility APIs power better computer use — dotey · 2026-09-15
- 5 production apps, 2M monthly requests for $6: why developers are going all-in on Cloudflare — viksit · 2026-09-15
- Mixing a 3090 with an Intel Arc B70 for 56GB VRAM local LLM inference? — overand · 2026-09-15
- Google Cloud and Inferact partner to make TPU a first-class vLLM target — vllm_project · 2026-09-15