New research traces distillation length inflation to student-teacher EOS token mismatch
tw_killian · x · 2026-09-20
New work from UNC, BYU, and Microsoft traces length inflation in on-policy distillation to student/teacher EOS mismatch: models can already solve the problem, then waste most of their token budget repeating output. The authors find that aligning decoding stop sets alone is not enough to fix the issue.
Related event: EOS Token Mismatch Behind Longer Distilled Outputs, Study Finds(2 posts)→
More from Research
- Sylvester's 1879 conjecture fully proved after 147 years — IgorCarron · 2026-09-20
- XGEN Labs unveils generative world simulation JING+DAO, tops WBench leaderboard — hey_abusiddik · 2026-09-20
- Schmidhuber: LLMs aren't truly creative because they lack compression progress — SchmidhuberAI · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- François Fleuret nicknames his training curves; researchers admit they curse baselines too — giffmana · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20