Cross-tokenizer on-policy distillation: supervision reliability beats alignment coverage
Bingxi Hou · hf · 2026-10-07
A new paper examines on-policy distillation across heterogeneous tokenizers: strict 1:1 aligned groups already cover most student-generated tokens, and restricting reverse KL to a student-selected top-16 shared-vocabulary subset matches full shared-vocabulary OPD and beats cross-tokenizer baselines. Adding MSE span supervision on mismatch groups surprisingly hurts accuracy, with weak or negative gradient agreement. The takeaway: prioritize supervision reliability over maximizing alignment coverage.
More from Research
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07