SimpleOPD Boosts Short-Context Reasoning via Distillation

Shanghai-AI-Laboratory · hf · 2026-08-17

Shanghai AI Laboratory proposes SimpleOPD, a tokenizer-agnostic on-policy distillation method. It distills knowledge from long-context reasoning teachers to short-context students, improving mathematical proof reasoning and generalizing to science benchmarks via token span alignment and length constraint.

Original post →

More from Research

Research channel →