New Handbook Breaks Down How Mixture of Experts Works, From Router Logits to Expert Parallelism
techNmak · x · 2026-09-15
techNmak published a handbook on how Mixture of Experts actually works, starting from what happens when a Transformer FFN becomes a routed bank of experts with only a small subset selected per token.
Topics covered: router logits, Top-k selection, expert weighting, load balancing, capacity and token dropping, router z-loss, dropless execution, total vs. active parameters, and expert parallelism.
Related event: Developer's MoE Handbook Goes Viral: From Routing to Expert Parallelism(3 posts)→
More from Research
- Sakana's PC-ALM trains 1000-layer nets without backprop; LeCun calls it target prop reborn — ylecun · 2026-09-15
- Five exercises find four defects in the author's own LLM-judge eval system — alexpran · 2026-09-15
- How a Major Children's Hospital Uses Open Source NVIDIA AI for Cardiac Care — nordicinst · 2026-09-15
- CMU-led preprint frames strategic motor learning as hypothesis testing — AnnaCiaunica · 2026-09-15
- Reasoning-Filtered Dataset Fable-5.1-Max-Filtered-5000x Trends on Hugging Face — MoreThought · 2026-09-15
- AISTATS 2027 heads to Montreal, paper submissions open September — qberthet · 2026-09-15