Nat Lambert argues distillation mostly matters in SFT and midtraining, not RL
natolambert · x · 2026-07-22
Nat Lambert argues that distillation matters most in SFT and midtraining, not as a magical trick inside RL.
- He says Anthropic has claimed some Chinese labs use its models in RL-shaped data pipelines, but that those samples are likely too small to drive major gains.
- His main point is that distillation is useful when it seeds reasoning behaviors during SFT and midtraining.
- He also argues that the way Chinese models say they are Claude-like reflects what next-token training actually learns: features and token clusters, not just surface imitation.
- In a reply, he rejects the idea that Chinese labs are using the strongest models as teachers during RL, saying that is not how distillation works and would not create such a large uplift.
Related event: Nat Lambert Clarifies LLM Distillation Mechanics(2 posts)→
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22