Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x decode speedup via MoE expert substitution

Zestyclose_Reality15 · reddit · 2026-10-06

The author optimized MoE offloading for Qwen3-Next-80B-A3B (Q4KM, 48.5GB) on an RTX 3090: only 1/4 of experts stay in VRAM, the rest stream from NVMe. Instead of always waiting for the router's expert, the patch substitutes in-VRAM experts when the router's pick misses — waiting only for the top pick and gate weight ≥0.15 experts cuts the quality cost from +5.7% ppl (full substitution) to +1.2% in the real engine. Decode throughput at 15.7GB VRAM jumps from stock llama.cpp's 72.9/65.8/31.8 tok/s (plenty/32GB/16GB RAM) to 108.4/97.8/94.5 tok/s; 89 tok/s even reading every miss from SSD at 16GB. Cost: +1.6% ppl, −1.8 GSM8K points, no HumanEval difference. Caveats: Qwen3-Next only, Linux+CUDA, single sequence, and the 16GB case was RAM-locked simulation. Code, scripts, logs and a paper (including what didn't work) are public.

Original post →

More from Infra

Infra channel →