Kimi K3 uses QAT, RL and MOPD across a wide expert-task mix
nrehiew_ · x · 2026-07-29
The thread outlines Kimi K3’s training pipeline and task mix.
- The pipeline is described as SFT with QAT using MXFP4 weights and MXFP8 activations, followed by RL and then MOPD.
- They train many expert models, including one for each reasoning level.
- The poster notes that the RL algorithm appears to be a GRPO variant rather than GLM 5.2 PPO.
- MOPD is used to consolidate domain-specialized capabilities, but top-k distillation did not show a clear advantage.
- The task categories include general verifiable tasks, kernel optimization, personal assistant tasks, autoresearch, and web development tasks, with SVG/3D explicitly mentioned.
More from Models
- Macaron-V1-Tall trends on Hugging Face as a text-generation model — mindlab-research · 2026-07-29
- Kimi's Open Model Pricing Sparks Debate: The $20M Monetization Reality — BenBajarin · 2026-07-29
- User Finds Claude Opus Overly Verbose, Switches to Sonnet for Better Focus — brandon_galang · 2026-07-29
- Meta paper says RL can optimize code speed, with Qwen 2.5 7B and CWM 32B gains — burny_tech · 2026-07-29
- User reverses course and says GPT 5.6 Sol is actually a really good model — TheZachMueller · 2026-07-29
- Sam Altman teases GPT-5.6 Sol on Cerebras at 750 tokens/sec in July — daniel_mac8 · 2026-07-29