Kimi K3 Architecture: Scaling Context, Depth Beyond Bigger MoE

AndLukyane · x · 2026-08-04

This article provides an in-depth analysis of the Kimi K3 model architecture by Moonshot AI. K3 scales both pre-training and post-training simultaneously, reaching 2.8T total parameters (104B activated per token) and incorporating 1M-token agentic trajectories.

The architecture organizes scaling around three types of information flow:

Original post →

More from Models

Models channel →