MoE Paper Explained: Sparsely-Gated Experts Boost Model Capacity Over 1000x

goyalshaliniuk · x · 2026-10-02

Part 5 of an AI-concepts-via-papers thread explains Mixture-of-Experts: instead of activating the whole model per input, routing sends inputs to different expert components, covering sparse computation, expert routing, scaling capacity, and compute efficiency.

The linked paper is the 2017 classic by Noam Shazeer, Geoffrey Hinton, Jeff Dean et al., which introduced a sparsely-gated MoE layer of up to thousands of feed-forward sub-networks. A trainable gating network picks a sparse expert combination per example, achieving >1000x capacity gains with minor efficiency losses on language modeling and machine translation benchmarks, with a 137B-parameter MoE beating prior state of the art at lower compute.

Original post →

More from Research

Research channel →