ik_llama.cpp MoE fork: expert residency and hybrid execution give 20-30% speedup on 6GB GPUs

IceFog72 · reddit · 2026-10-06

Author IceFog72 ships a fork of ikllama.cpp adding bandwidth-adaptive CPU/GPU execution, shared-expert residency, n-gram cache, and Q20 support, with a configurable VRAM cap via --moe-resident auto --moe-resident-mib N. On an RTX 2060 6GB running Qwen3.6-35B-A3B Q4, throughput rose from 23 t/s to 26-30 t/s (20-30%). It only helps when the full MoE doesn't fit in VRAM and the GPU has idle compute; the post includes full commands and tuning advice.

Original post →

More from Infra

Infra channel →