llama.cpp mmap fits Qwen3.8-Flash-Next in 16G+64G RAM at 26t/s
q8019222 · reddit · 2026-08-28
A Reddit user reports running Qwen3.8-Flash-Next (IQ3XSS quant) on a 16GB RAM + 64GB swap machine using llama.cpp's mmap, still hitting 26 tokens/s — faster than a non-MoE 30B model at 10 t/s on the same setup. mmap loads weights on demand, a practical trick for running large MoE models on modest hardware.
More from Infra
- Qualcomm's data center opportunity becomes tangible with custom ARM CPUs — BenBajarin · 2026-08-28
- Merge API Gateway Sees 11x Month-Over-Month Spike in Token Volume — shensi · 2026-08-28
- Local Inference Tool ds4 Adds Support for GLM 5.3 Flash — lakySK · 2026-08-28
- Running Qwen3.8-Flash on RTX 3090: Quantization & Performance — crusaderky · 2026-08-28
- Gemini 1.5 Flash inference speeds may exceed 300 tok/sec — Sentdex · 2026-08-28
- Paper analyzes NVSHMEM: System-level insights into GPU communication — thoefler · 2026-08-28