Kimi K3 Architecture: Scaling Context, Depth Beyond Bigger MoE
AndLukyane · x · 2026-08-04
This article provides an in-depth analysis of the Kimi K3 model architecture by Moonshot AI. K3 scales both pre-training and post-training simultaneously, reaching 2.8T total parameters (104B activated per token) and incorporating 1M-token agentic trajectories.
The architecture organizes scaling around three types of information flow:
- Sequence Length: Interleaves 3 efficient Kimi Delta Attention (KDA) layers with 1 global Gated MLA layer to balance cost and information retention.
- Depth: Introduces Attention Residuals, allowing layers to selectively retrieve representations from earlier blocks.
- Width: Utilizes Stable LatentMoE, expanding the model to 896 routed experts (activating 16 per token).
More from Models
- Databricks says Kimi K3 now runs at 239 tokens per second on its serving stack — Yuchenj_UW · 2026-08-04
- Verifying GPT Pro: Zero Math Errors, But Disastrous Exposition — josh_wills · 2026-08-04
- Ant Group Engineer's Long Post: The Four Very Different Bets of Chinese AI Labs — AcanthisittaOk1699 · 2026-08-04
- OpenAI's Upcoming 'Astra' Model to Focus on Multi-Agent Collaboration — thesaraharminta · 2026-08-04
- Testing Qwen3.8-Max: Open Models Are Catching Up with Closed Frontier — dair_ai · 2026-08-04
- Kimi K3 Estimated at 2.8T Params, Potentially Distilled from Smaller Opus — gabriberton · 2026-08-04