RL is the story: arguing a new model beats its synthetic data source after training

JoshPurtell · x · 2026-09-03

Josh Purtell disputes claims that a new model was distilled from sol-synthesized data: synthesizing via sol offers no advantage over using Kimi k3 or GLM 5.3 directly (the latter are cheaper), and the model outperforming sol after training suggests RL is a major factor. He adds SFT+RL is harder than RL-only, making lazy distillation the only plausible motivation.

Related event: JoshPurtell breaks down the model distillation debate: task type determines risk and feasibility(8 posts)→

Original post →

More from Models

Models channel →