68GB Qwen MoE at 21 tok/s on RTX 3060 + 16GB RAM, bit-exact via new --moe-direct-io
zyxciss · reddit · 2026-10-10
A Redditor implemented --moe-direct-io in llama.cpp, running the 68GB Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) IQ2XS build at 20-21 tok/s (24+ warm) on an RTX 3060 12GB + 16GB DDR4 machine — a 10-15x speedup over stock llama.cpp's 1.4-2.1 tok/s.
Key details:
- Two months ago the author tried predicting which MoE experts fire next to speed up CPU/GPU offload, but the early version pruned cold experts, hurting quality, so it was shelved;
- Stock llama.cpp suffers 1,100-1,565 major page faults per token and pulls 200-300MB from SSD; alternatives like Strata need 24GiB of experts pinned via mlock, impossible with 16GB RAM;
- The new blocking+prefetch design cuts page faults to 0 (just +1 across 32 tokens) while staying bit-exact to stock output — no experts dropped.
A genuine breakthrough for running large MoE models on low-RAM consumer hardware.
More from Infra
- Inference Auctions paper: serving highest bidders first without wrecking latency optimizations — nhaghtal · 2026-10-10
- Inference Auctions: bidding for scarce LLM serving capacity without breaking KV-cache gains — nhaghtal · 2026-10-10
- SGLang-Diffusion serving framework for diffusion models to be unveiled at PyTorch Conference 2026 — PyTorch · 2026-10-10
- Andrew Ng: one of my agents makes 5,000–10,000 web searches a day, data centers will need to scale far beyond current plans — DeepLearningAI · 2026-10-10
- SGLang Summit 2026 lineup: Intel CEO, Lilian Weng, Perplexity CEO, Nov 12-13 in SF — ying11231 · 2026-10-10
- AI inference could hit $350B next year, pushing software margins from 72% down to 30-50% — Sethwinterroth · 2026-10-10