LSD Online Distillation Cuts RL Post-Training Length Tax from 19% to -3.7%
Xu Wan · hf · 2026-10-01
Researchers quantify the "length-scaling tax" (LST)—excess verbosity on already-solved queries during RL post-training—and propose Length Self-Distillation (LSD), routing solved prompts to on-policy distillation with an EMA teacher while keeping RL for unsolved ones.
- No external teacher model required
- LST drops from 19.0% to -3.7% on single-turn reasoning, and 31.4% to 13.7% on multi-turn agentic tasks
- LSD matches or beats RL performance while keeping easy-query responses concise
More from Research
- New Survey Unifies 4D Dynamic Scene Reconstruction Across NeRF and 3DGS Approaches — zhenjun_zhao · 2026-10-01
- LBDU-VIO Cuts Position Error by 25.1% on EuRoC During 10s Visual Outages via Learned Bias Dynamics — zhenjun_zhao · 2026-10-01
- A visual refresher on the basics of Markov chains — alexbilz · 2026-10-01
- A $1 million prize for scientific honesty could reshape research culture — skdh · 2026-10-01
- Agent0: zero-data self-evolving agent framework from Stanford/Salesforce headed to COLM2026 — yuyinzhou_cs · 2026-10-01
- Melting Pot updated: Lab2d ships modern Python wheel, no more sandboxed old versions — jzl86 · 2026-10-01