Microsoft Research’s ReOPD reuses teacher prefixes to distill multi-turn agents offline

MicrosoftResearch · hf · 2026-07-24

ReOPD turns multi-turn distillation into an offline prefix replay problem

Microsoft Research studies on-policy distillation (OPD) for multi-turn agent tasks and proposes Replayed-Prefix On-Policy Distillation (ReOPD).

The core idea is to convert expensive agent-environment interaction into a reusable offline resource for scalable distillation.

Original post →

More from coding & agent

coding & agent channel →