AWS paper proposes SCOUT to fix off-policy mismatch in on-policy distillation

AWS · hf · 2026-10-01

A paper from AWS identifies a fundamental asymmetry in on-policy distillation (OPD): trajectories sampled by the student are off-policy for the teacher, whose continuation quality degrades as student-generated prefixes grow longer.

They propose SCOUT (Student-COnditioned Updates of the Teacher), a co-training framework that periodically adapts the teacher to student prefixes using RL with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards.

Controlled experiments confirm the mechanism, and across multiple teacher-student configurations, model scales, and reasoning domains, SCOUT consistently improves OPD effectiveness.

Original post →

More from Research

Research channel →