Kimi K3 paper details SiTU-GLU and quantile balancing for 896-expert MoE
KyeGomezB · x · 2026-07-28
A Kimi K3 paper excerpt lays out three MoE engineering changes meant to keep very sparse routing stable at scale:
- SiTU-GLU replaces the usual SwiGLU-style gating to reduce exploding activations in large models.
- Quantile balancing is used instead of standard auxiliary-loss load balancing when routing roughly 1,000 experts, with histogram-based bias updates to keep expert usage perfectly balanced.
- A LatentMoE variant moves RMSNorm to after expert aggregation, instead of the usual pre-norm placement.
The excerpt frames these changes as necessary for scaling to 896 routed experts with 16 active experts per token.
Related event: Kimi K3 Architecture: Introducing SiTU-GLU and Full-Rank Projection(3 posts)→
More from Models
- Kimi K3 report adds NoPE, full-rank gating and FP32 attention training — suchenzang · 2026-07-28
- Kimi K3 lands on Fireworks AI for inference and training — omarsar0 · 2026-07-28
- Gemini video generation adds words and blocks some harmless prompts — Individual-Cookie615 · 2026-07-28
- Polymarket now prices a 42% chance of a new Claude Sonnet by next month — Polymarket · 2026-07-28
- vLLM Collaborates with DigitalOcean to Host Kimi K3 Model — vllm_project · 2026-07-28
- Kimi K3 lands on ChatLLM with U.S. hosting and an open-source fine-tune — bindureddy · 2026-07-28