Cross-tokenizer on-policy distillation: supervision reliability beats alignment coverage

Bingxi Hou · hf · 2026-10-07

A new paper examines on-policy distillation across heterogeneous tokenizers: strict 1:1 aligned groups already cover most student-generated tokens, and restricting reverse KL to a student-selected top-16 shared-vocabulary subset matches full shared-vocabulary OPD and beats cross-tokenizer baselines. Adding MSE span supervision on mismatch groups surprisingly hurts accuracy, with weak or negative gradient agreement. The takeaway: prioritize supervision reliability over maximizing alignment coverage.

Original post →

More from Research

Research channel →