MoE expert expansion ported to llama.cpp: route more experts than native top-K at runtime, no training
Specific-Tax-6700 · reddit · 2026-09-07
A developer proposed a training-free inference architecture for sparse MoE models and ported expert expansion to llama.cpp: running MoE models with more routed experts than the native top-K (8→x), with an adaptive threshold, 99→50% linear influence decay, and per-layer range control—aiming to increase active parameters for more succinct reasoning without any training or fine-tuning.
The change is runtime-only and works across all backends. Tested on Qwen 3.6 35B A4B+; docs and code are on the moe-expansion branch of a llama.cpp fork on GitHub.
Related event: Dev builds llama.cpp fork enabling MoE expert expansion(2 posts)→
More from Infra
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 — himanshustwts · 2026-09-07
- Wan2GP lands on Pinokio: one-click AI video generation for 6GB+ VRAM machines — cocktailpeanut · 2026-09-07
- Wan2GP AMD edition hits Pinokio, supporting all RDNA 2-4 discrete GPUs — cocktailpeanut · 2026-09-07
- Buying a room full of hardware to run OpenClaw as supreme rage bait — HankYeomans · 2026-09-07
- Google says high-performance memory now exceeds 75% of an AI server's bill of materials — Beth_Kindig · 2026-09-07
- Fully automated product demo videos: local LLMs, 34 episodes, zero human editing — Ok_Cartographer_6086 · 2026-09-07