Moonshot’s Kimi K3 is a 2.8T-parameter MoE model with only 16 experts active

PeterDiamandis · x · 2026-07-29

Peter Diamandis highlights Moonshot AI’s Kimi K3 as a 2.8-trillion-parameter model that activates only 16 of 896 experts per token.

The post uses that architecture to make a broader point about sparsity: intelligence does not necessarily require every parameter to be used on every computation. In other words, very large models can still behave efficiently if only a small subset of experts is routed for each token.

Related event: Kimi K3 shifts attention from scale to architecture(25 posts)→

Original post →

More from Models

Models channel →