Kimi K3 report claims 2.5x scaling efficiency, 104B active params and 1M context
burny_tech · x · 2026-07-28
- Moonshot’s Kimi K3 technical report says the model scales three axes together: long context, depth, and width.
- It uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, claiming about 2.5× efficiency over Kimi K2.
- The model has 104B activated parameters out of 2.8T total, does not use RoPE, and instead relies on NoPE + recurrent decay gates to reach 1M tokens of context without retuning positional embeddings.
- For vision, the team trains the encoder from scratch with next-token prediction and reports more stability than starting from SigLIP while matching vision evals.
- Post-training includes Multi-Teacher On-Policy Distillation (MOPD), merging 9 RL-trained specialist teachers into one unified model.
Related event: Kimi K3 Tech Report: 2.8T MoE and Architectural Efficiency Breakthroughs(49 posts)→
More from Models
- Testing K3 and Open Models: Reasoning Tokens Can Burn Entire Budgets — MaziyarPanahi · 2026-07-28
- Local Test: Nanbeige4.2-3B Lags Behind Qwen MoE in KV Cache Efficiency — TechTefa · 2026-07-28
- 35B Agentic Bakeoff: KAT-Coder Matches Qwen at Half the Token Cost — IvGranite · 2026-07-28
- LangChain Event: Open Models Match Closed Frontier in Agent Tasks — LangChain · 2026-07-28
- Tabular foundation models aim to replace LLMs on structured data — bendee983 · 2026-07-28
- Dev Team Drops Claude Opus for Coding, Keeps It Only for Specific Tasks — MicahBerkley · 2026-07-28