Running Qwen3.8 Flash Next on 128GB RAM + one 5080: 136pp/19tg at Q5_K_XL

whatyathinkk · reddit · 2026-09-15

A detailed local-inference setup for Qwen3.8 Flash Next on 128GB RAM + a single RTX 5080 (16GB): Unsloth's Q5KXL hits 136pp/19tg, Q4KXL 152pp/23tg, Q3KXL 195pp/23tg. Slow but impressive quality, the poster says.

Config highlights: 6-part GGUF (UD-Q5KXL-ncmoe48), lazy-mode on, ngl=999 with n-cpu-moe=48 keeping MoE layers on CPU, perlayertokenembd forced to CPU, 262K context, q80 KV cache, flash-attn on; sampling at temp=1.0/top-k=20/top-p=0.95 with thinking preserved and reasoningeffort=medium. Full llama.cpp config included; the poster asks about MTP or other speedups.

Original post →

More from Infra

Infra channel →