Moonshot’s Kimi K3 report says its largest model has 2.78T parameters and 2.5× better scaling
teortaxesTex · x · 2026-07-27
- Moonshot’s Kimi K3 technical report says the model is its largest ever, with 2.78T total parameters, 104.2B activated parameters, and 2.5× better scaling efficiency than Kimi K2.
- The report describes a native multimodal training recipe, a hybrid KDA-MLA attention setup, and training choices such as per-head Muon, weight clipping, QB load balancing, cosine decay, and 1% warmup.
- It also says Kimi K3 starts at 8k context and is later extended to 64k, while the model can extrapolate to 1M-token contexts without RoPE rescaling or interpolation.
- The accompanying Hugging Face page shows moonshotai/Kimi-K3 live on HF, with the model card, eval results, and deployment metadata visible.
Related event: Moonshot Open-Sources Kimi K3: A 2.8T Parameter Multimodal Model(19 posts)→
More from Models
- Kimi K3 reportedly improves training efficiency by 2.5× — zephyr_z9 · 2026-07-27
- Kimi K3 goes live on Baseten Model APIs at day 0 — baseten · 2026-07-27
- NVIDIA distills Cosmos3 Super image-to-video to 4 steps with a 64B model — multimodalart · 2026-07-27
- Kimi K3 goes live on Modal with custom DFlash speculative decoding for lossless speedup — AAAzzam · 2026-07-27
- Kimi K3 with 2.8T parameters and 1M context now supported on vLLM — ricklamers · 2026-07-27
- Kimi K3 could become a cheap distillation base for personal, local models — victormustar · 2026-07-27