Nat Lambert says RL distillation does not use the strongest models as teachers
natolambert · x · 2026-07-21
Nat Lambert pushes back on a claim about Chinese labs using the strongest models as teachers during RL distillation.
- He argues that this is not how distillation works.
- Using the strongest model as a teacher during RL would not deliver a huge lift, because graders in RL are messy and expensive.
- The thread frames the debate as a practical training/evaluation question, not a matter of simple model copying.
More from Research
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22
- DepthART scales monocular depth to tiny models, hitting 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22
- Project CETI gets a Jeopardy! shout-out with a SETI-style whale clue — begusgasper · 2026-07-22