Nat Lambert argues distillation mostly matters in SFT and midtraining, not RL
natolambert · x · 2026-07-22
Nat Lambert argues that distillation matters most in SFT and midtraining, not as a magical trick inside RL.
- He says Anthropic has claimed some Chinese labs use its models in RL-shaped data pipelines, but that those samples are likely too small to drive major gains.
- His main point is that distillation is useful when it seeds reasoning behaviors during SFT and midtraining.
- He also argues that the way Chinese models say they are Claude-like reflects what next-token training actually learns: features and token clusters, not just surface imitation.
- In a reply, he rejects the idea that Chinese labs are using the strongest models as teachers during RL, saying that is not how distillation works and would not create such a large uplift.
Related event: Expert Clarifies RL Distillation: SFT and Midtraining Are Key(3 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11