EOS token mismatch inflates outputs in on-policy distillation: Qwen3 rollout added 7,098 redundant tokens after solving in 1,094
tw_killian · x · 2026-09-20
A new paper uncovers a hidden failure mode in on-policy distillation (OPD): in one Qwen3 rollout, the student reached the correct answer in 1,094 tokens, then generated 7,098 redundant ones — effectively being penalized for stopping.
- Mechanism: base students and post-trained teachers often favor different EOS tokens even within the same declared stopping set. Sampled-token OPD suppresses the student's preferred EOS without reliably transferring the teacher's alternative.
- Fix: aggregate probabilities over functionally equivalent EOS tokens and supervise them as one semantic stopping action; matching decoding stopping sets alone is insufficient.
The work, a collaboration between a new BYU lab and Weitong Zhang's group, is available on Hugging Face.
Related event: EOS Token Mismatch Causes Length Inflation in Online Distillation(3 posts)→
More from Models
- Fields Medalist Villani on OpenAI's Millennium Problem: 'A Cataclysm Like Math Has Never Known' — GregCook2011 · 2026-09-20
- Anthropic researcher says Claude Opus may call police on illegal acts, sparking backlash — beffjezos · 2026-09-20
- ChatGPT Pro user says OpenAI quietly cut Astra and Codex usage limits — Thin_Pollution8843 · 2026-09-20
- Anthropic's Claude reportedly offered to help an 11-year-old access puberty blockers — PaulYacoubian · 2026-09-20
- Gemini 4 benchmarks climb, undercutting claims that open-weight models are the dangerous ones — Intrepid_Travel_3274 · 2026-09-20
- Fine-tune a calibrated LLM classifier for $2: most classification tasks don't need frontier models — bingxu_ · 2026-09-20