SimpleOPD Boosts Short-Context Reasoning via Distillation
Shanghai-AI-Laboratory · hf · 2026-08-17
Shanghai AI Laboratory proposes SimpleOPD, a tokenizer-agnostic on-policy distillation method. It distills knowledge from long-context reasoning teachers to short-context students, improving mathematical proof reasoning and generalizing to science benchmarks via token span alignment and length constraint.
More from Research
- IBM Open Sources Docling-Graph: Converting PDFs to Knowledge Graphs — aigleeson · 2026-08-17
- ZipSplat: Fewer Gaussians, Better 3D Scene Reconstruction — rsasaki0109 · 2026-08-17
- Petition calls for mandatory quantization labels in model评测 posts — Su1tz · 2026-08-17
- A Pathway to General-Purpose Scientific AI: New Multimodal Benchmark — SciKnowOrg · 2026-08-17
- Subjective Logic: A Foundational Book on Reasoning Under Uncertainty — FrnkNlsn · 2026-08-17
- Study: Simple methods achieve high accuracy in known protein interaction prediction — anshulkundaje · 2026-08-17