Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context
alex_verem · x · 2026-07-29
Moonshot’s Kimi K3 technical report argues that the next leap in AI may come from using compute more efficiently, not just scaling it up.
What the model is
- Kimi K3 is a 2.8T-parameter open-weight MoE model with 104B activated parameters.
- It supports a 1-million-token context window.
- The architecture includes Kimi Delta Attention and Attention Residuals to improve information flow across sequence length and depth.
- Its Stable LatentMoE design activates only 16 of 896 experts per token, with the goal of routing compute where it matters instead of firing the whole model.
Why the team says it matters
The report claims roughly 2.5× better intelligence per unit of compute versus Kimi K2.
Reported results
Post-training reinforcement learning across general, coding, and reasoning tasks is said to improve compositional generalization and long-horizon execution.
At 2.8T scale, the system is presented as a combined effort in algorithm design, KDA-balanced expert-parallel training, memory management, million-token agentic RL, persistent rollout, sandbox states, and deployment innovations.
Benchmarks and positioning
- The report says Kimi K3 reaches frontier-level results in long-horizon coding, agentic tasks, knowledge, reasoning, and vision.
- It also claims the model trails only the strongest proprietary systems in their evaluated suite.
- Moonshot released the full weights to support future research and adoption.
More from Infra
- Extropic signs a $75 million Commerce Department LOI to scale thermodynamic AI chips — beffjezos · 2026-07-29
- SK hynix reportedly signs long-term contracts with about 10 major customers — dejavucoder · 2026-07-29
- Shanghai Aishengna is said to be manufacturing DUV lithography tools — zephyr_z9 · 2026-07-29
- A vendor-agnostic Vulkan backend cuts edge inference latency from 30 ms to 3 ms — ppchaos · 2026-07-29
- pdf-mcp turns technical PDFs into structured text, images, and searchable context — tom_doerr · 2026-07-29
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29