DuoOPD uses joint teacher-student outcomes for on-policy distillation, +5.98 points over OPD

Ao Yu · hf · 2026-10-01

On-policy distillation pushes down even the student's correct responses because it ignores success outcomes. DuoOPD lets the student's outcome set the feedback direction and the joint teacher-student outcome decide the teacher's support: the teacher's verified answer scores failed student responses when only the teacher succeeds, and a task-shared weight reinforces whole responses when only the student succeeds.

One rule covers all four outcome combinations. Across Qwen3 and Llama, DuoOPD beats all five baselines on mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and also leads on task mixtures spanning scientific calculation, instruction following, and code generation.

Original post →

More from Research

Research channel →