EOS token mismatch inflates outputs in on-policy distillation: Qwen3 rollout added 7,098 redundant tokens after solving in 1,094

tw_killian · x · 2026-09-20

A new paper uncovers a hidden failure mode in on-policy distillation (OPD): in one Qwen3 rollout, the student reached the correct answer in 1,094 tokens, then generated 7,098 redundant ones — effectively being penalized for stopping.

The work, a collaboration between a new BYU lab and Weitong Zhang's group, is available on Hugging Face.

Related event: EOS Token Mismatch Causes Length Inflation in Online Distillation(3 posts)→

Original post →

More from Models

Models channel →