Running Qwen3.8-Flash-Next on a 5090 with llama.cpp: 40 tok/s and barely any RAM used

nirurin · reddit · 2026-09-25

A first-hand report of running Qwen3.8-Flash-Next (Atomic Q4KM quant, 33 GGUF shards) via llama.cpp on a 5090 + 64GB RAM. Using mmap/lazy-mode flags with --n-cpu-moe 32, the setup uses 27GB VRAM and only 8GB system RAM at 40 tok/s decode and 50 tok/s prompt processing — surprisingly fast, with parts apparently streaming from NVMe rather than being buffered in RAM.

Original post →

More from Infra

Infra channel →