New curriculum traces on-policy distillation from Hinton's soft targets to frontier post-training
burkov · x · 2026-09-04
Andrey Burkov released a new ChapterPal curriculum, "Must-read papers on on-policy distillation":
- Traces knowledge distillation from Hinton's original soft-target objective to multi-teacher on-policy pipelines used in frontier post-training today
- Key shift: training a student on a fixed corpus of teacher outputs vs. sampling from the student itself — the change that addressed exposure bias and reshaped how LMs are compressed and aligned
- Shows the technique applied at scale in technical reports of today's largest models
Readable with an AI tutor on ChapterPal.
More from Research
- Mol-JEPA: A Multimodal JEPA Foundation Model for Molecules, One Year in the Making — TerribleAntelope9348 · 2026-09-04
- Two Years On: Six Guidelines for Making Research Impact via Open-Source in AI — lateinteraction · 2026-09-04
- GPT-6 Astra claims SOTA on ARC-AGI-3 at 66%, up from Sol's 8% — teortaxesTex · 2026-09-04
- Steering Qwen along a grader-vs-human dimension oddly shifts its personality — voooooogel · 2026-09-04
- CMU's AI Reviewer Beats Best Human Reviewer, Featured by Science — AkariAsai · 2026-09-04
- Chollet: ARC-AGI-4 lands Q1 2027, and solving ARC-3 is not AGI — fchollet · 2026-09-04