SimpleOPD: Solving Token Mismatch in Long-Context Distillation
bronzeagepapi · x · 2026-08-18
A new paper introduces SimpleOPD, a tokenizer-agnostic on-policy distillation method to transfer reasoning capabilities from long-context teachers to short-context students. Addressing issues like tokenizer mismatch and length explosion, the method operates in a shared text space and aligns identical spans. Experiments on models like Qwen3 and GLM-4.7 show consistent gains in mathematical reasoning.
Related event: Shanghai AI Lab Proposes SimpleOPD for Online Policy Distillation(2 posts)→
More from Research
- Exploring generative Gabor wavelets: a novel approach to non-photorealistic image synthesis — pixlpa · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24
- Claude Verifies 43 Lean Modules autonomously, Tackling Theoretical Physics — Tkaraletsos · 2026-08-24