Kimi K3 report says a 2.8T MoE used RL experts and multi-teacher distillation
cwolferesearch · x · 2026-07-29
A detailed breakdown of Kimi K3’s post-training recipe says the model is a 2.8T-parameter open-weight MoE with 104B activated parameters. The report describes a multi-stage pipeline: cold-start SFT with synthetic agentic trajectories and human-in-the-loop verification, per-domain RL across different effort levels, and then multi-teacher on-policy distillation to fuse nine expert teachers into one unified model.
The post argues that this recipe is becoming the standard pattern for frontier post-training: train narrow experts with RL, then consolidate them through distillation. It also notes that the report does not describe a separate RLHF or DPO stage, and claims Kimi K3 delivers around 2.5× overall scaling efficiency over Kimi K2, with strong results in coding, general reasoning, and long-horizon agentic tasks.
More from Models
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29
- Kimi K3 Tech Report: How Moonshot Achieved 2.5x Compute Efficiency — alex_verem · 2026-07-29
- Experienced web developer says Claude Opus 4.8 and 5 are now unusable for chat — FuzzyHead455 · 2026-07-29
- Reddit user says paid Gemini Pro access is still routing to other models — IAmMonke2 · 2026-07-29
- A fake Claude “model welfare” leak turns into an AI-community meme — repligate · 2026-07-29
- ChatGPT Android app appears to replace model picker with effort levels — justanothergin · 2026-07-29