RL is the story: arguing a new model beats its synthetic data source after training
JoshPurtell · x · 2026-09-03
Josh Purtell disputes claims that a new model was distilled from sol-synthesized data: synthesizing via sol offers no advantage over using Kimi k3 or GLM 5.3 directly (the latter are cheaper), and the model outperforming sol after training suggests RL is a major factor. He adds SFT+RL is harder than RL-only, making lazy distillation the only plausible motivation.
More from Models
- Gemini 3.8 Flash Lands: Third Flash Update in 6 Weeks, Boosting Agentic and Coding Skills — brianryhuang · 2026-09-03
- Startup Mostik bridges AI models via their weights, tops ARC-AGI 3 at 1/20 the cost — nordicinst · 2026-09-03
- Anthropic launches browser tool to detect Claude-made files via C2PA content credentials — btibor91 · 2026-09-03
- Marin 535B A23B Training: Blog and WandB Metrics Now Public — Sentdex · 2026-09-03
- ByteDance's looped language models match 12B rivals at 1.4B size, with Bengio as co-author — peterjliu · 2026-09-03
- Marin 535B A23B Frontier-Scale Training Run Is Fully Livestreamed — Sentdex · 2026-09-03