Distillation debate: RL, not distilling from sol, likely explains the model's gains
JoshPurtell · x · 2026-09-03
JoshPurtell pushes back on claims that a certain model distilled from sol:
- Synthesizing training data with sol offers almost no benefit over Kimi k3 or GLM 5.3 — the latter are actually cheaper
- Key evidence: the model outperforms sol after training, so RL is the big part of the story
- If they distilled, they wouldn't add an SFT stage to the RL pipeline (SFT+RL is harder than RL-only), so distillation is possible but "not 80%"
For long-horizon agent developers distillation is tempting; for trivial one-shot structured-output tasks with gold outputs, synthesizing data is trivially easy.
More from Models
- Leak: Astra won't be the best model of the year; a 'monster' is slated for end of year — ChrisGPT · 2026-09-03
- Startup Mostik bridges AI models via their weights, tops ARC-AGI 3 at 1/20 the cost — nordicinst · 2026-09-03
- Anthropic launches browser tool to detect Claude-made files via C2PA content credentials — btibor91 · 2026-09-03
- ByteDance's looped language models match 12B rivals at 1.4B size, with Bengio as co-author — peterjliu · 2026-09-03
- Marin 535B A23B Frontier-Scale Training Run Is Fully Livestreamed — Sentdex · 2026-09-03
- Gemini 3.8 Flash builds a working game from a single prompt, no frontier model needed — ColbyHawker · 2026-09-03