Moonshot opens Kimi K3 with 104B activated parameters and a 1M-token context window

teortaxesTex · x · 2026-07-27

Moonshot says Kimi K3 is a 2.8T MoE model with 104B activated parameters, native visual understanding, and a 1M-token context window.

The post also links the weight release and technical report, and the attached screenshot shows the model summary: 93 layers, 896 experts, 16 experts selected per token, and a hybrid attention stack with KDA and gated MLA. Moonshot frames the architecture as delivering “2.5x the intelligence per unit of compute,” and is opening up more of the stack, including attention kernels, MoE communication, and agent-environment infrastructure.

Related event: Moonshot Open-Sources Kimi K3: 2.8T Params and $20M Commercial Threshold(148 posts)→

Original post →

More from Models

Models channel →