Study: privileged references add little to on-policy self-distillation beyond distillation itself
NationalUniversityofSingapore · hf · 2026-09-18
This paper isolates what privileged references actually add to on-policy self-distillation (OPSD). Using AMPLE-Math—5,319 math problems with six reasoning views sharing the same answer—the authors show reference-free distillation accounts for most of Qwen3-1.7B's gains; extra reference benefit is modest in Qwen and worth only 2 points in SmolLM3-3B at step 50. Gains depend on the student actively training, and swapping short direct-response rollouts for thinking-enabled ones turns gains into losses. OPSD largely improves cross-mode transfer between direct-answer and thinking inference; references matter only insofar as they aid that transfer.
More from Research
- Turning Yang-Mills existence and mass gap into a formal conjecture is AI math's ultimate test — geoffreyirving · 2026-09-18
- CoRL 2026 SPIN workshop pits robots against a child in live dexterity challenge — berkeley_ai · 2026-09-18
- That viral "air-gapped data exfiltration" paper only read temperature over a 4cm gap — basedjensen · 2026-09-18
- An AI forecaster has won the seasonal Metaculus Cup for the first time — NathanpmYoung · 2026-09-18
- New video: weather forecasting with neural networks explained — ariG23498 · 2026-09-18
- First LLM runs entirely on-chain: weights, activations, attention all in SVM transactions — AccBalanced · 2026-09-18