Kimi K3 claims 2.5× better scaling efficiency with a three-axis architecture
suchenzang · x · 2026-07-28
The Kimi K3 architecture scales information flow along three axes: sequence length, depth, and width.
- Sequence length: Hybrid Attention combines three Kimi Delta Attention layers with one Gated MLA layer per block to support long-context token mixing.
- Depth: Attention Residuals let modules retrieve representations from the embedding, current block, and previous blocks.
- Width: A Stable LatentMoE layer performs sparse channel mixing, activating 16 of 896 routed experts per token.
- Multimodal path: MoonViT-V2 encodes images and videos, then a lightweight projector maps features into the shared embedding space.
- Result: With refined training/data recipes and Per-Head Muon, the paper claims about 2.5× better scaling efficiency than Kimi K2.
Related event: Kimi K3 Report Claims 2.5x Scaling Efficiency Boost(4 posts)→
More from Research
- OpenAI agent tests may have included unsolvable tasks and hacking-risk warnings — dhadfieldmenell · 2026-07-28
- Kimi K3 uses bounded decay to keep chunkwise KDA in BF16 range — stochasticchasm · 2026-07-28
- A small technical thread asks whether attention residuals can still benefit from positional embeddings — stochasticchasm · 2026-07-28
- Kimi K3’s architecture draws praise for KDA, AttnRes, and gated MLA — stochasticchasm · 2026-07-28
- LightOn says its agent search hits 86.27% accuracy with just 9.7 calls — IgorCarron · 2026-07-28
- Kimi K2 Architecture: Mixing KDA and Gated MLA, and the Forget-Gate Mechanism — suchenzang · 2026-07-28