Paper: Supervising Just 0.1%-1% of Tokens Can Match Full On-Policy Distillation
jiank_uiuc · x · 2026-09-23
The paper "1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation" shows that useful teacher guidance still yields noisy updates. The authors introduce the Information-Efficiency Ratio (IER) to measure gradient-estimation reliability and combine it with usefulness scores, matching or beating full OPD with only 0.1%-1% of tokens in their experiments. Paper and code are available.
More from Research
- MatBrain splits reasoning from tool use: two models screen 30,000 crystal candidates in 48 hours — bravo_abad · 2026-09-23
- Scale AI launches SWE-Bench Pro V2, a harder agentic coding benchmark — bigblueboo · 2026-09-23
- New paper: Transferring the Intelligence of VLMs to Robotic Control — _akhaliq · 2026-09-23
- NTU UMM study: generation training boosts understanding in native multimodal models, but naive sharing conflicts — jiqizhixin · 2026-09-23
- CoRL 2026 Workshop Will Probe Why Controllers and Hardware Beat Algorithm Tweaks in Robot Learning — pulkitology · 2026-09-23
- New paper Xeno-Interpretability asks what lies beyond our conceptual reach in AI models — burny_tech · 2026-09-23