Moonshot’s Kimi K3 report says sparse MoE and delta attention beat bigger dense models
SumitGup · x · 2026-07-26
Moonshot AI’s Kimi K3 technical report argues that bigger dense models are not the future, and instead highlights five design ideas: sparse routing, delta attention, attention residuals, a 1M-token context window, and efficient scaling.
- Sparse MoE routing: 896 experts exist, but only 16 are activated per token.
- Delta attention: focuses computation on what changed instead of recomputing full context.
- Attention residuals: preserve useful signal across layers.
- 1M context: aims to fit entire repositories, books, and papers in one session.
- Efficient scaling: targets frontier coding performance without frontier inference costs.
Related event: Moonshot AI Unveils Trillion-Parameter Model Kimi K3(2 posts)→
More from Models
- Gary Marcus says video understanding remains unreliable outside training-like inputs — GaryMarcus · 2026-07-26
- Google Search is now showing 5+ ads per page as AI Overviews may hit monetization — mobileraj · 2026-07-26
- Flux 3 impresses users with coherent dialogue from extremely vague video prompts — cocktailpeanut · 2026-07-26
- SemiAnalysis argues Google’s TPU deals may be holding Gemini back — firstadopter · 2026-07-25
- GPT 5.6 Sol builds a working 3D game from just a few prompts — Angaisb_ · 2026-07-25
- Google’s Gemma open-weight models are being praised for strong industrial fine-tuning — clmt · 2026-07-25