Kimi K3 lands as a 2.8T MoE with 1M tokens and open weights
moonshotai · hf · 2026-07-28
- Moonshot releases Kimi K3, a 2.8T-parameter MoE model with 104B activated parameters, native vision, and a 1M-token context window.
- The model introduces Kimi Delta Attention and Attention Residuals to improve information flow, plus Stable LatentMoE that activates 16 of 896 experts per token.
- The company says these changes deliver roughly 2.5× better scaling efficiency than Kimi K2, with post-training focused on general, agentic, and coding RL across multiple reasoning-effort levels.
- Kimi claims frontier-level results on long-horizon coding, agentic, knowledge, reasoning, and vision tasks, while still trailing the strongest proprietary models in its suite; the full weights are released.
Related event: Moonshot Releases Kimi K3 Open Weights: 2.8T Parameters Sets New Record(137 posts)→
More from coding & agent
- Users say Grok’s setup is smoother than Codex and needs less extra tooling — yangyi · 2026-07-28
- MAPD distills agentic search into structured protocols and lifts Qwen3 scores — Junlin Liu · 2026-07-28
- JarvisHub turns a canvas into shared memory for multimodal creative agents — Yunlong Lin · 2026-07-28
- Codex and Claude Code browsers are being used to run mini apps as live agent tools — RileyRalmuto · 2026-07-28
- Practitioner says multi-agent workflows work better with 1–2 active agents, then long autonomous runs — edgarpavlovsky · 2026-07-28
- Only one of 63 API companies exposes all three agent-ready surfaces — ExtensionPea834 · 2026-07-28