Kimi K3 uses Stable LatentMoE to tame exploding activations in a 896-expert design
suchenzang · x · 2026-07-28
Kimi K3’s Stable LatentMoE design targets two failure modes in large MoE layers: exploding internal activations and unstable load balancing.
What changes
- Adds RMSNorm before the up-projection.
- Replaces the usual gate with SiTU-GLU to keep activations bounded.
- Uses Quantile Balancing (QB) for routing load balance.
Why it matters
The image explains that Kimi K3 scales channel mixing to 896 routed experts with 16 active experts per token. The paper argues that the extreme sparsity makes the vanilla MoE design numerically fragile, and the new components are meant to stabilize both the routed branch and the expert assignment process.
Related event: Inside Kimi K3: Tri-axis Architecture and Hybrid Attention(35 posts)→
More from Models
- Moonshot releases Kimi K3, a 2.8T open-weight multimodal agent with 1M context — johnseach · 2026-07-28
- Moonshot opens Kimi K3 weights, a 2.8T MoE model with 1M-token context — chris_j_paxton · 2026-07-28
- Kimi K3 goes live on OpenRouter as third-party providers race to add support — scaling01 · 2026-07-28
- Kimi K3 arrives on Applied Compute for training and inference — rhythmrg · 2026-07-28
- A repost claims Anthropic’s Opus 5 regresses badly despite benchmark gains — rickasaurus · 2026-07-28
- Polymarket prices a 76% chance Moonshot ships another Kimi K model by September — Polymarket · 2026-07-28