A handbook on how Mixture of Experts actually works, from routing to expert parallelism
techNmak · x · 2026-09-15
The author compiled a comprehensive MoE handbook covering router logits, Top-k selection, expert weighting, load balancing, capacity and token dropping, router z-loss, dropless execution, total vs. active parameters, and expert parallelism — with worked examples referencing Mixtral, DeepSeek-V3, and Qwen3, all cited back to original papers and current implementations.
Related event: Developer's MoE Handbook Goes Viral: From Routing to Expert Parallelism(3 posts)→
More from Research
- Microsoft's ESRL boosts MoE RL via expert-space exploration — MicrosoftResearch · 2026-09-15
- Grouped Value Attention shrinks KV cache by reconstructing keys on demand — Vishesh Tripathi · 2026-09-15
- Amazon's MInTRL uses sparse off-policy interventions to boost on-policy RL — amazon · 2026-09-15
- Stateless LLM failover preserves ~0% context; ContinuityBench proxy hits 99.20% CPR — its_vayishu · 2026-09-15
- Phillip Isola highlights a non-mainstream AI route: RL from scratch via ultra-fast simulators — AjdDavison · 2026-09-15
- Cutting AI verifier reading cost: top-50 retrieval kept just 2 of 8 minority evidence items — iMiguelmars · 2026-09-15