AlignOPSD fixes decision-timestamp mismatch in agent distillation, beating GRPO by 5.5-8.7%

Mingju Chen · hf · 2026-09-29

Researchers propose AlignOPSD to fix a "decision-timestamp mismatch" in on-policy self-distillation for long-horizon agents: privileged teacher guidance often lands on the wrong timestep, and a student's decision may span multiple steps, so timestamp-local supervision misaligns both context and credit scope.

The method works in two stages:

Evaluated on ALFWorld, WebShop, and Search-QA with Qwen2.5-3B/7B, AlignOPSD beats GRPO and StepOPSD in all eight backbone-metric comparisons, improving over GRPO by 5.5-8.7% and ranking first in six. Code is open-sourced.

Original post →

More from coding & agent

coding & agent channel →