Kimi K3 technical report shows 2.5× better scaling efficiency than Kimi K2
stochasticchasm · x · 2026-07-28
The image is from the Kimi K3 technical report and highlights two main points:
- Scaling efficiency: Kimi K3 is shown as achieving a 2.5× gain in scaling efficiency over Kimi K2.
- Architecture changes: compared with Kimi K2, Kimi K3 increases total parameters from 1.04T to 2.78T, activated parameters from 32.6B to 104.2B, and training context length from 128K to 1M.
The table also shows a larger MoE setup, more routed experts, more attention heads, and a new hybrid KDA–MLA attention mechanism.
Related event: Inside Kimi K3: Tri-axis Architecture and Hybrid Attention(34 posts)→
More from Research
- Different coding agents can produce very different estimates from the same model spec — JessicaHullman · 2026-07-28
- CMU’s O-VAD detects industrial video anomalies by tracking object state over time — CarnegieMellonU · 2026-07-28
- Latent Action Model talk will show how to learn Super Mario from observation alone — ceciletamura · 2026-07-28
- GLM 5 follows up with per-head Muon to balance attention-head updates — stochasticchasm · 2026-07-28
- MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16 — zephyr_z9 · 2026-07-28
- OpenAI chart says 43.5% of occupation-specific ChatGPT use goes beyond the user’s job — soumitrashukla9 · 2026-07-28