New paper: supervising just 1% of tokens can match full on-policy distillation, 0.1% sometimes suffices
jiank_uiuc · x · 2026-09-23
A new MBZUAI–Ant Group paper, "1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation," studies why useful teacher guidance in sparse on-policy distillation (OPD) can still yield noisy updates when gradients are estimated from a single sampled token.
Key points:
- Introduces the Information-Efficiency Ratio (IER), derived from a signal-to-noise decomposition, to measure gradient-estimation reliability at a given prefix
- IER can be combined with existing usefulness scores for token selection while keeping the sampled reverse-KL objective
- On math and medical reasoning tasks, 0.1%–1% token budgets match or beat full OPD without selection
- Gains hold in both thinking-on and thinking-off modes, strongest at small budgets
Paper and code are public.
More from Models
- Anthropic unveils Claude Opus 5.5, days after Dario Amodei's call to "pace the frontier" — Polymarket · 2026-09-23
- Grok is now in your Tesla: hands-free inbox, calendar and file work via Connectors — yunta_tsai · 2026-09-23
- Claude Code 2.1.280 ships 114 changes, makes Opus 5.5 with 1M context the default — ClaudeCodeLog · 2026-09-23
- Claude Code 2.1.280 ships 114 CLI changes, defaults to Opus 5.5 at $4/$20 per Mtok — ClaudeCodeLog · 2026-09-23
- Anthropic Teases Claude Sonnet 5.5 and Haiku 5.5 Launching in Coming Weeks — mikeyk · 2026-09-23
- Opus 5.5 Lauded as One-Shot Machine, 25% Cheaper Than Opus 5 — bindureddy · 2026-09-23