Moonshot opens Kimi K3 with 104B activated parameters and a 1M-token context window
teortaxesTex · x · 2026-07-27
Moonshot says Kimi K3 is a 2.8T MoE model with 104B activated parameters, native visual understanding, and a 1M-token context window.
The post also links the weight release and technical report, and the attached screenshot shows the model summary: 93 layers, 896 experts, 16 experts selected per token, and a hybrid attention stack with KDA and gated MLA. Moonshot frames the architecture as delivering “2.5x the intelligence per unit of compute,” and is opening up more of the stack, including attention kernels, MoE communication, and agent-environment infrastructure.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11