Distillation or frontier training? AI circle debates the real source of model gains

xeophon · x · 2026-09-25

A debate over where frontier model gains actually come from: one side argues the improvements are "largely" from distillation (and thus SFT), questioning whether the approach bottoms out—"some model" must actually be able to perform the task, or you have to buy trajectories from labellers.

xeophon pushes back, noting that argument doesn't hold: you can target different pass rates during training, and to advance the frontier a model has to be trained in a way that genuinely explores, not just distills existing capability. The exchange touches the core tension between distillation and self-exploration in RL for LLMs.

Related event: AI Researchers Debate Whether Distillation Drives China's Model Gains(5 posts)→

Original post →

More from Models

Models channel →