Tsinghua shows on-policy distillation works with a single training example

jiqizhixin · x · 2026-09-24

Tsinghua University's new paper, Rethinking On-Policy Distillation II: One Training Example, runs an extreme experiment: shrinking an OPD training set from 17,000 problems to just one.

The team notes data scale/composition effects on OPD remain under-explored versus training objectives and optimization.

Original post →

More from Research

Research channel →