Moonshot’s Kimi K3 report says its largest model has 2.78T parameters and 2.5× better scaling
teortaxesTex · x · 2026-07-27
- Moonshot’s Kimi K3 technical report says the model is its largest ever, with 2.78T total parameters, 104.2B activated parameters, and 2.5× better scaling efficiency than Kimi K2.
- The report describes a native multimodal training recipe, a hybrid KDA-MLA attention setup, and training choices such as per-head Muon, weight clipping, QB load balancing, cosine decay, and 1% warmup.
- It also says Kimi K3 starts at 8k context and is later extended to 64k, while the model can extrapolate to 1M-token contexts without RoPE rescaling or interpolation.
- The accompanying Hugging Face page shows moonshotai/Kimi-K3 live on HF, with the model card, eval results, and deployment metadata visible.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23