Kimi K3 paper describes a three-stage post-training stack with SFT, RL, and MOPD
stochasticchasm · x · 2026-07-28
Kimi K3 uses SFT, RL, and MOPD to merge specialized agent policies
The post shows a paper excerpt describing Kimi K3’s post-training pipeline as a three-stage process:
- Supervised fine-tuning (SFT) to build a cold-start policy.
- Reinforcement learning (RL) to develop domain-specific experts at different reasoning levels.
- Multi-Teacher On-Policy Distillation (MOPD) to consolidate those specialized policies into one model.
The screenshot adds that the SFT dataset was expanded for complex agentic tasks using prior Kimi-series models, followed by multi-stage verification and human-in-the-loop annotation.
Related event: Kimi K3 Technical Report Details 3-Stage Post-Training(4 posts)→
More from coding & agent
- Moonshot AI and Together AI to explain how Kimi K3 powers production agent workflows — togethercompute · 2026-07-28
- Kimi K3 launches with 2.8T parameters, 1M context and $0.30 input pricing — togethercompute · 2026-07-28
- Together AI brings Moonshot’s Kimi K3 online with 1M context and agent tools — togethercompute · 2026-07-28
- Part III benchmarks Opus 5 on SlopCodeBench in the software-factory debate — teropa · 2026-07-28
- The key to this prompt is a sub-agent blind A/B critic — mattshumer_ · 2026-07-28
- A single prompt and Opus 5 generated a Three.js game workflow — majidmanzarpour · 2026-07-28