Sparse distillation paper: supervising just 0.1%-1% of tokens can match or beat full OPD
jiank_uiuc · x · 2026-09-23
A new arXiv paper, 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation, tackles a hidden inefficiency in sparse on-policy distillation (OPD): even useful teacher guidance produces noisy updates when the gradient is estimated from a single sampled next token.
Method
- Working in information geometry, the authors derive an information-efficiency ratio (IER) from a signal-to-noise decomposition under an optimal scalar baseline — higher IER means lower relative gradient estimation error.
- IER can be approximated with a small candidate set and combined with existing usefulness selectors via IER-OR / IER-AND, keeping the sampled reverse-KL objective unchanged.
Results
- On math reasoning with a 0.1% token budget (JustRL-Qwen3-4B → Qwen3-1.7B, thinking off, Bayes@32), IER-based selection beats full OPD: AIME25 14.4→19.4, AIME26 12.5→16.3, HMMT25 8.7→11.7, HMMT26 13.4→14.4.
- On HealthBench, roughly one supervised token per trajectory matches full OPD.
- Gains hold in both thinking-on and thinking-off modes, and larger token budgets sometimes add little or hurt — which tokens you supervise matters more than how many.
Paper and code are available.
More from Models
- Anthropic Teases Claude Sonnet 5.5 and Haiku 5.5 Launching in Coming Weeks — mikeyk · 2026-09-23
- Opus 5.5 Lauded as One-Shot Machine, 25% Cheaper Than Opus 5 — bindureddy · 2026-09-23
- No single model to rule them all: developer juggles Opus, GPT-6 Astra and Chinese models as agents — iannuttall · 2026-09-23
- Opus 5.5 Jumps to 1st Place on the AA Intelligence Index — UnknownEssence · 2026-09-23
- Bizarre Benchmark Result Prompts Advice to Stick With Opus 5.5 Medium — Yuchenj_UW · 2026-09-23
- OpenAI models scale normally, Claude has "earthquake-shaped" scaling, says Lambert — natolambert · 2026-09-23