SFT+RL pipeline choice makes the lazy-distillation theory unlikely, argues dev
JoshPurtell · x · 2026-09-03
Follow-up in JoshPurtell's distillation debate thread:
- If they lazily distilled from sol, the only rationale would be not wanting to spin up Kimi on openrouter — but that doesn't hold either
- Distilling would mean not adding an SFT stage to the RL pipeline (especially with good starting performance), and SFT+RL is harder than RL-only
- Conclusion: possible but not 80%
Complements the earlier post: the model's gains likely come from RL, not distillation.
More from Models
- Dev: hard to trust frontier labs' data promises, another reason to use OSS models — adityaag · 2026-09-03
- NVIDIA tops Hugging Face open-source repos with 500+ added, ahead of Alibaba and Tencent — Sam_Witteveen · 2026-09-03
- GLM 5.3 at 200+ TPS builds a full web page in 37 seconds, unedited video — nutlope · 2026-09-03
- Same-prompt test: fable 5.1 vs sol 5.6 generating a realistic three.js waterfall — majidmanzarpour · 2026-09-03
- Claude Fable 5.1 tops Code Arena with 1765 pts, 77-pt lead over Qwen — burny_tech · 2026-09-03
- New Google model debuts at #14 and #32 on leaderboards, mocked vs Kimi and GLM — airesearch12 · 2026-09-03