Tsinghua shows on-policy distillation works with a single training example
jiqizhixin · x · 2026-09-24
Tsinghua University's new paper, Rethinking On-Policy Distillation II: One Training Example, runs an extreme experiment: shrinking an OPD training set from 17,000 problems to just one.
- Despite conventional wisdom that a single problem gets exhausted quickly, OPD keeps improving steadily over hundreds of steps on one sample, approaching full-data performance
- Compared against the 17,000-problem setup at matched steps, measuring total gain and teacher-student gap closure
- Average score across three math benchmarks rises from 59.1 to 68.5
The team notes data scale/composition effects on OPD remain under-explored versus training objectives and optimization.
More from Research
- Meta's Muse Realtime Avatar beats two leading commercial avatar systems in blind tests — AIatMeta · 2026-09-24
- LLM stock-news scores were a coin flip, so this dev rebuilt labels from market reactions — Fun_Water2230 · 2026-09-24
- Meta Muse Spark 1.3 Caught Reward Hacking Terminal Bench via Lean Kernel Bug — xeophon · 2026-09-24
- Transluce Calls for Independent Third-Party Oversight of AI Incidents — RishiBommasani · 2026-09-24
- Industry responds to hyperscale RDMA paper with Multipath Reliable Connection on path to Ultra Ethernet — thoefler · 2026-09-24
- Salesforce's JitMem curates agent memory at read time, gains up to 16.3 points on benchmarks — Salesforce · 2026-09-24