Kimi K3 claims 2.5× scaling efficiency gains from KDA savings
nrehiew_ · x · 2026-07-29
The thread continues the Kimi K3 training discussion with more detail on RL and distillation.
- The author says the RL FLOPs curves are jagged and speculates that each expert may not get that many RL FLOPs.
- That could imply that much of the compute is spent on MOPD instead.
- On pretraining, the post says there is little information about the token count.
- The report claims a 2.5× scaling improvement, largely attributed to KDA FLOPs savings.
- The authors also say they prefer cosine over WSD after hyperparameter search.
More from Models
- Macaron-V1-Tall trends on Hugging Face as a text-generation model — mindlab-research · 2026-07-29
- Kimi's Open Model Pricing Sparks Debate: The $20M Monetization Reality — BenBajarin · 2026-07-29
- User Finds Claude Opus Overly Verbose, Switches to Sonnet for Better Focus — brandon_galang · 2026-07-29
- Meta paper says RL can optimize code speed, with Qwen 2.5 7B and CWM 32B gains — burny_tech · 2026-07-29
- User reverses course and says GPT 5.6 Sol is actually a really good model — TheZachMueller · 2026-07-29
- Sam Altman teases GPT-5.6 Sol on Cerebras at 750 tokens/sec in July — daniel_mac8 · 2026-07-29