Shanghai AI Lab Proposes SimpleOPD for Online Policy Distillation
Shanghai AI Lab introduces SimpleOPD, a tokenizer-agnostic online policy distillation method that transfers reasoning from long-context teachers to short-context students, improving math and science reasoning while addressing token mismatch and response length explosion.
2026-08-17 ~ 2026-08-18 · 2 related posts
- SimpleOPD Boosts Short-Context Reasoning via Distillation — Shanghai-AI-Laboratory · 2026-08-17
- SimpleOPD: Solving Token Mismatch in Long-Context Distillation — bronzeagepapi · 2026-08-18