DuoOPD uses joint teacher-student outcomes for on-policy distillation, +5.98 points over OPD
Ao Yu · hf · 2026-10-01
On-policy distillation pushes down even the student's correct responses because it ignores success outcomes. DuoOPD lets the student's outcome set the feedback direction and the joint teacher-student outcome decide the teacher's support: the teacher's verified answer scores failed student responses when only the teacher succeeds, and a task-shared weight reinforces whole responses when only the student succeeds.
One rule covers all four outcome combinations. Across Qwen3 and Llama, DuoOPD beats all five baselines on mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and also leads on task mixtures spanning scientific calculation, instruction following, and code generation.
More from Research
- ArchMap lands in Nature Genetics: code-free single-cell mapping onto reference atlases — burny_tech · 2026-10-01
- New Paper Argues Literary Tools Are Essential for Building Culturally Literate AI — begusgasper · 2026-10-01
- NVIDIA's Instant NuRec reconstructs a drivable 3DGS world from driving logs in ~1.5 seconds — rsasaki0109 · 2026-10-01
- KV-streams trains SWE agents 2x faster by preserving KV cache across compaction — burny_tech · 2026-10-01
- Researcher argues "ego" beats "persona" for describing LLM identity — repligate · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01