Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x decode speedup via MoE expert substitution
Zestyclose_Reality15 · reddit · 2026-10-06
The author optimized MoE offloading for Qwen3-Next-80B-A3B (Q4KM, 48.5GB) on an RTX 3090: only 1/4 of experts stay in VRAM, the rest stream from NVMe. Instead of always waiting for the router's expert, the patch substitutes in-VRAM experts when the router's pick misses — waiting only for the top pick and gate weight ≥0.15 experts cuts the quality cost from +5.7% ppl (full substitution) to +1.2% in the real engine. Decode throughput at 15.7GB VRAM jumps from stock llama.cpp's 72.9/65.8/31.8 tok/s (plenty/32GB/16GB RAM) to 108.4/97.8/94.5 tok/s; 89 tok/s even reading every miss from SSD at 16GB. Cost: +1.6% ppl, −1.8 GSM8K points, no HumanEval difference. Caveats: Qwen3-Next only, Linux+CUDA, single sequence, and the 16GB case was RAM-locked simulation. Code, scripts, logs and a paper (including what didn't work) are public.
More from Infra
- vLLM Semantic Router makes Mixture-of-Models programmable; evals show instructions swing agent MCP tool use from 0/40 to 20/40 — colinmcnamara · 2026-10-06
- NVIDIA survey: 89% of telcos say open models are key to their AI strategy — NVIDIAAI · 2026-10-06
- Dev's custom vLLM patch runs DiffusionGemma on DGX Spark, edging out the commercial API — vllm_project · 2026-10-06
- How a Solo Dev Stopped Local 7B Models from Hallucinating Bank Balances: Code Calculates, Model Summarizes — Revibed69 · 2026-10-06
- Cloudflare launches cf, an agent-first CLI to query observability data via the API — dinasaur_404 · 2026-10-06
- Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options — DrainBramage · 2026-10-06