Kimi K3 pairs a 2.8T MoE with 1M context and multi-teacher post-training
novasarc01 · x · 2026-07-28
Kimi K3 is a 2.8T MoE model with native vision and 1M-token context
The post summarizes the Kimi K3 technical report and its training recipe.
- Kimi K3 is described as a 2.8T-parameter MoE model with 104B activated parameters, native vision capability, and a 1-million-token context window.
- The report says it uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE to improve scaling efficiency by about 2.5× over Kimi K2.
- Post-training combines RL across general, agentic, and coding domains, plus multiple reasoning-effort levels.
- The model uses Multi-Teacher On-Policy Distillation (MOPD) to unify specialized behaviors into one deployable checkpoint.
- Evaluation claims frontier-level performance on long-horizon coding, agentic tasks, knowledge, reasoning, and vision, while still trailing Claude 4.5 and GPT-5.6 Sol overall.
Related event: Kimi K3 Report: Trillion-Parameter MoE and Post-Training Innovations(2 posts)→
More from coding & agent
- Netlify says 60% of June 2026 signups made their first deploy through Drop — thisiskp_ · 2026-07-28
- Emblem says it can build an 80-slide cited deck in 35 minutes instead of 100 hours — rohanpaul_ai · 2026-07-28
- Claude Opus 5 keeps running Python scripts to edit files in coding workflows — samuelcolvin · 2026-07-28
- Vercel highlights its open-source stack, including Eve and agent-browser — evilrabbit_ · 2026-07-28
- LightOn says its agent search hits 86.27% accuracy with just 9.7 calls — IgorCarron · 2026-07-28
- The Modern Wait Equation: Fix AI Slop Now or Wait for the Next Model? — ilanbigio · 2026-07-28