Paper pinpoints EOS token mismatch as the root failure mode of on-policy distillation
tw_killian · x · 2026-09-21
Sharing a paper that solves a known failure mode in on-policy distillation: student rollouts become excessively verbose and exhaust the generation budget. The mechanical culprit is termination-token mismatch between base students and post-trained teachers—even with identical declared stop sets, their probability mass sits on different EOS tokens. Every student stop attempt is treated as a wrong continuation, penalizing and suppressing the student's termination instinct without transferring the teacher's preferred alternative. Notably, aligning decoding stop sets alone doesn't fix it; the real fix treats functionally equivalent EOS tokens as equivalent.
More from Research
- Jev's eval abstraction maps 1:1 to autorubric paper from 8 months ago, researcher finds — deliprao · 2026-09-21
- Researcher Says TypeSafe's Jev Mirrors His Autorubric LLM Eval Framework From 8 Months Ago — deliprao · 2026-09-21
- As ICLR 2027 Tops 60K Submissions, Researcher Proposes 3-4 Paper Cap per Author — ziv_ravid · 2026-09-21
- ICLR 2027 Hits 60K+ Submissions; Researcher Proposes Paper Caps, Forced Reproducibility — ziv_ravid · 2026-09-21
- No, Laya isn't capped at 512 tokens — it's ModernBERT with 8192-token configs — antoine_chaffin · 2026-09-21
- Kev open-source decision models scale to 0.6B/4B/8B, trainable in 40 min on one H100 — TheMoonMidas · 2026-09-21