SlimWise prunes MoE experts only at decode, boosting throughput up to 1.81x
Gunho Park · hf · 2026-10-07
- Problem: MoE models activate few experts per token, but batched decoding touches nearly the whole expert pool, making expert-weight traffic a bottleneck; conventional pruning also prunes compute-bound prefill, hurting quality for little gain.
- Approach: SlimWise tailors the expert pool per phase — full model for prefill, pruned model for decode, reusing the prefill KV cache directly — plus a low-cost distillation stage so the decoder can continue from full-model KV caches.
- Deployment: implemented in vLLM, supporting both PD disaggregation and PD-colocated serving.
- Results: on Qwen3.6-35B-A3B, up to 1.81x decode throughput at 50% expert pruning with minimal accuracy loss; the paper also shows benchmark scores can hide pruning-induced generation-length distortions.
More from Infra
- Tracking LLM API model deprecations and rolling alias changes across providers — shamikhan005 · 2026-10-07
- One of the Last American Chestnut Groves to Be Destroyed for a Data Center — Promptmethus · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- The 2019 Mac Pro with 1.5TB RAM would be the ultimate local LLM machine today — Odd-Capital-847 · 2026-10-07
- Team claims sub-5-second full weight sync for 1T-parameter RL training — saurabh_shah2 · 2026-10-07
- NVIDIA's UNREAL: one model unifies corpus retrieval and long-context at 128K+ — nvidia · 2026-10-07