Moonshot releases Kimi K3, a 2.8T MoE model with a 1M-token context window

burny_tech · x · 2026-07-28

Kimi K3 debuts as a 2.8T MoE model with 1M-token context

Moonshot says Kimi K3 is its most capable model yet: a 2.8T MoE model with native visual understanding and a 1M-token context window. The company claims a new architecture that delivers 2.5x intelligence per unit of compute, framing the release as an efficiency leap rather than a pure parameter-scale race.

Alongside the model weights and technical report, Moonshot is also opening parts of the stack behind K3, including high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale. A comment in the thread notes the model uses aggressive sparsity: around 2% active experts, or 16 out of 896 experts total.

Related event: Moonshot Releases Open-Weight Kimi K3 Model(138 posts)→

Original post →

More from Infra

Infra channel →