Moonshot opens Kimi K3 with 104B activated parameters and a 1M-token context window
teortaxesTex · x · 2026-07-27
Moonshot says Kimi K3 is a 2.8T MoE model with 104B activated parameters, native visual understanding, and a 1M-token context window.
The post also links the weight release and technical report, and the attached screenshot shows the model summary: 93 layers, 896 experts, 16 experts selected per token, and a hybrid attention stack with KDA and gated MLA. Moonshot frames the architecture as delivering “2.5x the intelligence per unit of compute,” and is opening up more of the stack, including attention kernels, MoE communication, and agent-environment infrastructure.
Related event: Moonshot Open-Sources Kimi K3: 2.8T Params and $20M Commercial Threshold(148 posts)→
More from Models
- Reply says the next Codex release may arrive Tuesday, with a Cerebras GPT-5.6xHigh update soon — eyishazyer · 2026-07-28
- Kimi K3 weights landed hours before a live discussion on hybrid architectures and agent harnesses — hugobowne · 2026-07-28
- Commenter says Anthropic is only ahead by thin margins, not by a wide model lead — intellectronica · 2026-07-28
- Google Gemma 4 Vision token-budget demo starts trending on Hugging Face — google · 2026-07-28
- GPT 5.6 Ultra is too unstable for agent workflows, author says — tensorqt · 2026-07-28
- Firecrawl tops a SimpleQA search benchmark with 94.7%, ahead of Exa and Claude — Candid-Dog-775 · 2026-07-28