MoE Paper Explained: Sparsely-Gated Experts Boost Model Capacity Over 1000x
goyalshaliniuk · x · 2026-10-02
Part 5 of an AI-concepts-via-papers thread explains Mixture-of-Experts: instead of activating the whole model per input, routing sends inputs to different expert components, covering sparse computation, expert routing, scaling capacity, and compute efficiency.
The linked paper is the 2017 classic by Noam Shazeer, Geoffrey Hinton, Jeff Dean et al., which introduced a sparsely-gated MoE layer of up to thousands of feed-forward sub-networks. A trainable gating network picks a sparse expert combination per example, achieving >1000x capacity gains with minor efficiency losses on language modeling and machine translation benchmarks, with a 137B-parameter MoE beating prior state of the art at lower compute.
More from Research
- Memorizon trains streaming world models beyond context window with only 12% step-time overhead — MBZUAI-IFM · 2026-10-02
- Morgan Stanley's Parallel Power Tempering sampling rivals RL post-training without weight updates — morganstanley · 2026-10-02
- KaliBench: 8,504 pairs benchmark shows no open-weight LLM exceeds 42% on Kali Linux CLI tasks — RISys-Lab · 2026-10-02
- DataMagic: multi-agent system turns raw data into data videos, +83% quality, 79.7% faster — Yupeng Xie · 2026-10-02
- Six Coding Agents, One Repo: Isolated Runs All Broke, Chatting Agents All Passed — jokiruiz · 2026-10-02
- Open lab: does a cheap decision model keep parallel coding agents from colliding? — jokiruiz · 2026-10-02