llama.cpp mmap fits Qwen3.8-Flash-Next in 16G+64G RAM at 26t/s

q8019222 · reddit · 2026-08-28

A Reddit user reports running Qwen3.8-Flash-Next (IQ3XSS quant) on a 16GB RAM + 64GB swap machine using llama.cpp's mmap, still hitting 26 tokens/s — faster than a non-MoE 30B model at 10 t/s on the same setup. mmap loads weights on demand, a practical trick for running large MoE models on modest hardware.

Related event: llama.cpp Users Optimize Qwen3.8-Flash-Next on Low-Memory and Dual-GPU Setups(2 posts)→

Original post →

More from Infra

Infra channel →