Kimi K3 arrives as a 2.8T MoE with 104B active params and a 1M-token context

stochasticchasm · x · 2026-07-28

A reply to the Kimi K3 report notes that the tech report does not disclose training tokens or compute for either K3 or MoonVIT 2.

The attached abstract shows why the model is notable: Kimi K3 is a 2.8T-parameter MoE with 104B activated parameters, native vision support, and a 1M-token context window. The paper claims roughly a 2.5× scaling-efficiency improvement over Kimi K2, cites work on KDA, Stable LatentMoE, and long-horizon post-training, and says K3 reaches frontier-level results on long-context coding, agentic, knowledge, reasoning, and vision tasks.

It also says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, but beats other open and proprietary models in the authors’ evaluation suite. The full weights are released to support further research and adoption.

Original post →

More from Models

Models channel →