ik_llama.cpp MoE fork: expert residency and hybrid execution give 20-30% speedup on 6GB GPUs
IceFog72 · reddit · 2026-10-06
Author IceFog72 ships a fork of ikllama.cpp adding bandwidth-adaptive CPU/GPU execution, shared-expert residency, n-gram cache, and Q20 support, with a configurable VRAM cap via --moe-resident auto --moe-resident-mib N. On an RTX 2060 6GB running Qwen3.6-35B-A3B Q4, throughput rose from 23 t/s to 26-30 t/s (20-30%). It only helps when the full MoE doesn't fit in VRAM and the GPU has idle compute; the post includes full commands and tuning advice.
More from Infra
- ~90% of frontier lab compute now goes to post-training and inference — IanAndrewsDC · 2026-10-06
- Qwen 27B on 2× RX 7900 XT: 66.5 TPS single-stream, still short of claimed 100+ — EqualCryptographer67 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06
- One Dot burns 1.6B tokens/day on a $100 subscription — roughly $540k/month in API-equivalent compute — DarthSilent · 2026-10-06
- Charles Frye (Modal) explains inference engines: schedulers, KV cache, CUDA graphs, speculative decoding — AI Engineer · 2026-10-06
- Reflection Ships Apache 2.0 Model With Tech Report; Analyst Estimates Pre-training MFU at Just ~12% — eliebakouch · 2026-10-06