Kimi K3 report says a 2.8T MoE used RL experts and multi-teacher distillation

cwolferesearch · x · 2026-07-29

A detailed breakdown of Kimi K3’s post-training recipe says the model is a 2.8T-parameter open-weight MoE with 104B activated parameters. The report describes a multi-stage pipeline: cold-start SFT with synthetic agentic trajectories and human-in-the-loop verification, per-domain RL across different effort levels, and then multi-teacher on-policy distillation to fuse nine expert teachers into one unified model.

The post argues that this recipe is becoming the standard pattern for frontier post-training: train narrow experts with RL, then consolidate them through distillation. It also notes that the report does not describe a separate RLHF or DPO stage, and claims Kimi K3 delivers around 2.5× overall scaling efficiency over Kimi K2, with strong results in coding, general reasoning, and long-horizon agentic tasks.

Related event: Deep Dive into Kimi K3 Architecture: 2.8T Parameters and Attention Innovations(16 posts)→

Original post →

More from Models

Models channel →