Running Qwen3.8-Flash-Next on a 5090 with llama.cpp: 40 tok/s and barely any RAM used
nirurin · reddit · 2026-09-25
A first-hand report of running Qwen3.8-Flash-Next (Atomic Q4KM quant, 33 GGUF shards) via llama.cpp on a 5090 + 64GB RAM. Using mmap/lazy-mode flags with --n-cpu-moe 32, the setup uses 27GB VRAM and only 8GB system RAM at 40 tok/s decode and 50 tok/s prompt processing — surprisingly fast, with parts apparently streaming from NVMe rather than being buffered in RAM.
More from Infra
- Oracle's Massive 'Project Jupiter' Data Center Declares Force Majeure, Jeopardizing US AI Rollout — AIFlow_ML · 2026-09-25
- Wasmer runs a real PostgreSQL 18.4 server on iOS and in the browser via WebAssembly — jedisct1 · 2026-09-25
- Inference startup Jatevo returns: 124B tokens, 1.66M requests, $147K of inference delivered — toptickcrypto · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Agentic AI changes the CPU-to-GPU ratio: 5% CPU allocation cuts token cost ~3.7% — BenBajarin · 2026-09-25
- PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving — aryaman2020 · 2026-09-25