Kimi K3 Architecture Analyzed: 3T Parameters with Sparser MoE, Compute Details Undisclosed

stochasticchasm · x · 2026-07-28

A developer analyzed the architecture of Moonshot's newly open-sourced Kimi K3 model. The model features 3T total parameters and adopts a sparser Mixture-of-Experts (MoE) configuration (16/896) compared to K2's 8/384.

The author notes the current sparsity is solid but could probably go even sparser, guessing that the bottleneck might be VRAM. Furthermore, they point out an interesting choice by the team: the technical reports for both Kimi K3 and Moonvit 2 do not disclose the training tokens or compute used anywhere.

Related event: Kimi K3 Architecture: 3T Parameters and Ultra-Sparse MoE(2 posts)→

Original post →

More from Models

Models channel →