Kimi K3 report shows 2.5× better scaling efficiency than Kimi K2
ricklamers · x · 2026-07-28
A post on Kimi K3 highlights a new architecture and a reported efficiency jump over Kimi K2.
- The comparison table shows Kimi K3 growing from K2’s 61 layers / 1.04T total parameters / 32.6B activated parameters to 93 layers / 2.78T total parameters / 104.2B activated parameters.
- Kimi K3 also extends training context length from 128K to 1M and changes the attention stack to a hybrid KDA–MLA design.
- The scaling plot claims 2.5× better scaling efficiency versus Kimi K2 in terms of validation loss at a given FLOP budget.
More from Models
- Kimi K3 reveals its pre-training mix across text, code, math, knowledge, and vision — stochasticchasm · 2026-07-28
- Kimi K3 deployment estimate points to 2 B200 nodes or 1 B300 class node — HarveenChadha · 2026-07-28
- YC interview with Claude Code creator Boris Cherny covers Opus 5, prompt injection, and 80% shorter system prompts — ycombinator · 2026-07-28
- Kimi K3 is weights-available, not open source, under a non-commercial license — juntao · 2026-07-28
- The Verge says Moonshot’s open Kimi K3 could undercut closed U.S. AI models — The Verge AI · 2026-07-28
- Kimi K3 lands on Fireworks AI for inference and training — omarsar0 · 2026-07-28