AlignOPSD fixes decision-timestamp mismatch in agent distillation, beating GRPO by 5.5-8.7%
Mingju Chen · hf · 2026-09-29
Researchers propose AlignOPSD to fix a "decision-timestamp mismatch" in on-policy self-distillation for long-horizon agents: privileged teacher guidance often lands on the wrong timestep, and a student's decision may span multiple steps, so timestamp-local supervision misaligns both context and credit scope.
The method works in two stages:
- Decision-Aligned Supervision Rectification: re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence.
- Semi-Markov Hierarchical Credit Assignment: derives variable-duration decision spans from correspondence changes and allocates outcome-grounded credit across spans and turns.
Evaluated on ALFWorld, WebShop, and Search-QA with Qwen2.5-3B/7B, AlignOPSD beats GRPO and StepOPSD in all eight backbone-metric comparisons, improving over GRPO by 5.5-8.7% and ranking first in six. Code is open-sourced.
More from coding & agent
- Community-maintained list tracks 51 active AI angel investors with sources and verification dates — TheMoonMidas · 2026-09-29
- 400+ LLM Agents Living in a 2004-era MMO Server, All Local on Qwen 4B — kristiantalley679 · 2026-09-29
- px0 release fixes LSP memory leaks, ReDoS vulnerability, adds CSV table rendering — arpit_bhayani · 2026-09-29
- Full Prompt Shared: Make Your Agent Build an Offline Progress Dashboard Before Long Tasks — myLifeintheStack · 2026-09-29
- Local gamedev AI stack: Meta's Muse Glimmer nails tool use in one try where Gemma 4 12b fails — draginol · 2026-09-29
- Founder: AI Can Now Do the Work of 10 Employees Before You Hire — FinanceYF5 · 2026-09-29