Distillation helps, but SFT and midtraining still do the heavy lifting
danielrock · x · 2026-07-22
A repost of a thread arguing that distillation is helpful but not a shortcut to a frontier model. The quoted discussion says the real leverage comes in SFT and midtraining, where reasoning behavior can be seeded more effectively than by simply copying trajectories.
It also notes Anthropic’s claim that DeepSeek and others may have used its models in an RL-shaped data pipeline, but that the sample counts were small and likely only useful for initialization or limited experiments. The broader point is that “they say they’re Claude” is not evidence of a direct fast path to frontier performance.
Related event: Expert Clarifies RL Distillation: SFT and Midtraining Are Key(3 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11