Distillation helps, but SFT and midtraining still do the heavy lifting

danielrock · x · 2026-07-22

A repost of a thread arguing that distillation is helpful but not a shortcut to a frontier model. The quoted discussion says the real leverage comes in SFT and midtraining, where reasoning behavior can be seeded more effectively than by simply copying trajectories.

It also notes Anthropic’s claim that DeepSeek and others may have used its models in an RL-shaped data pipeline, but that the sample counts were small and likely only useful for initialization or limited experiments. The broader point is that “they say they’re Claude” is not evidence of a direct fast path to frontier performance.

Related event: Expert Clarifies RL Distillation: SFT and Midtraining Are Key(3 posts)→

Original post →

More from Models

Models channel →