One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
OVHaiLLM · hf · 2026-09-08
A new study critically reviews on-policy self-distillation for mathematical reasoning.
- The method uses privileged information to guide a model without a larger teacher
- It suffers from reasoning collapse
- The paper identifies three governing factors: token weighting, privileged context, and guidance dynamics — "one symptom, three levers"
More from Research
- Terence Tao blogs on Buckmaster and Alpöge's blow-up solutions to fluid equations — LucaAmb · 2026-09-08
- ECCV 2026 to Host Interactive Social Avatars Workshop With Keynotes From Andrew Zisserman and More — JonathonLuiten · 2026-09-08
- Building an active monitor for a black-box LLM: where to draw the automation line — Particular-Roof4257 · 2026-09-08
- NYU mathematician says LLMs helped crack fluid blowup problems in a month, accuses OpenAI of scoop and authorship pressure — dotey · 2026-09-08
- EmbodiedSkills: A Unified Framework for Orchestrating VLA Robot Agents — Wei Wang · 2026-09-08
- Fluid Regularity Landscape Moving Fast: Tao Says Smooth NS Now Looks Feasible — basedjensen · 2026-09-08