Moonshot opens Kimi K3 with 104B activated parameters and a 1M-token context window

teortaxesTex · x · 2026-07-27

Moonshot says Kimi K3 is a 2.8T MoE model with 104B activated parameters, native visual understanding, and a 1M-token context window.

The post also links the weight release and technical report, and the attached screenshot shows the model summary: 93 layers, 896 experts, 16 experts selected per token, and a hybrid attention stack with KDA and gated MLA. Moonshot frames the architecture as delivering “2.5x the intelligence per unit of compute,” and is opening up more of the stack, including attention kernels, MoE communication, and agent-environment infrastructure.

Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→

Original post →

More from Models

Models channel →