SimpleOPD: Solving Token Mismatch in Long-Context Distillation

bronzeagepapi · x · 2026-08-18

A new paper introduces SimpleOPD, a tokenizer-agnostic on-policy distillation method to transfer reasoning capabilities from long-context teachers to short-context students. Addressing issues like tokenizer mismatch and length explosion, the method operates in a shared text space and aligns identical spans. Experiments on models like Qwen3 and GLM-4.7 show consistent gains in mathematical reasoning.

Related event: Shanghai AI Lab Proposes SimpleOPD for Online Policy Distillation(2 posts)→

Original post →

More from Research

Research channel →