AWS paper proposes SCOUT to fix off-policy mismatch in on-policy distillation
AWS · hf · 2026-10-01
A paper from AWS identifies a fundamental asymmetry in on-policy distillation (OPD): trajectories sampled by the student are off-policy for the teacher, whose continuation quality degrades as student-generated prefixes grow longer.
They propose SCOUT (Student-COnditioned Updates of the Teacher), a co-training framework that periodically adapts the teacher to student prefixes using RL with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards.
Controlled experiments confirm the mechanism, and across multiple teacher-student configurations, model scales, and reasoning domains, SCOUT consistently improves OPD effectiveness.
More from Research
- TraceML compares 4,465 human vs 207 agent Kaggle trajectories: agents use narrower research skills — burny_tech · 2026-10-01
- ST-AudioLM From Sony AI and POSTECH Tracks Moving Sound Sources With 40 Trajectory Tokens — mittu1204 · 2026-10-01
- Stanford Qi Lab Wins NIH TRDNT Challenge Phase I With Spatial RNA Reprogramming Platform — anshulkundaje · 2026-10-01
- WashU Prof Wins NIH R35 to Decode Human-Specific Gene Regulation with AI — anshulkundaje · 2026-10-01
- Factorizing attention matrices rewires gradient flow, destroying an information-exponent bottleneck — burny_tech · 2026-10-01
- Tsinghua and Stanford locate a reward subsystem in LLMs with value and dopamine neurons — jiqizhixin · 2026-10-01