1% of Tokens Can Be Enough: MBZUAI and Ant Group Improve On-Policy Distillation
jiank_uiuc · x · 2026-09-23
A MBZUAI–Ant Group paper (arXiv:2609.24432) studies gradient estimation in sparse on-policy distillation (OPD), where teacher supervision is allocated to a small subset of tokens:
- Problem: gradients estimated from sampled next tokens are noisy, so useful teacher guidance can yield noisy updates.
- Method: analyzing estimation at a fixed prefix via information geometry, the authors propose an information-efficiency ratio (IER) from a signal-to-noise decomposition, enabling token selection via a candidate-set approximation that composes with existing usefulness scores while keeping the sampled reverse-KL objective.
- Results: IER improves existing selectors on math and medical reasoning; at 0.1%–1% token budgets, sparse configurations match or exceed full OPD without selection. Gains hold in both thinking-on and thinking-off modes. Paper and code are public.
More from Research
- MatBrain splits reasoning from tool use: two models screen 30,000 crystal candidates in 48 hours — bravo_abad · 2026-09-23
- Scale AI launches SWE-Bench Pro V2, a harder agentic coding benchmark — bigblueboo · 2026-09-23
- If AI Writes All the Papers, Peer Review Becomes Humanity's Remaining Role — sudoraohacker · 2026-09-23
- Yarin Gal: I Ignore Papers Where the Candidate Isn't First or Last Author — yaringal · 2026-09-23
- New paper: Transferring the Intelligence of VLMs to Robotic Control — _akhaliq · 2026-09-23
- NTU UMM study: generation training boosts understanding in native multimodal models, but naive sharing conflicts — jiqizhixin · 2026-09-23