AgentOPSD: Recursive self-distillation for agentic RL achieves 89.1% on ALFWorld

Tsinghua · hf · 2026-08-07

Tsinghua proposes AgentOPSD, a critic-free recursive method for turn-level credit assignment in agentic RL. It aggregates token-level teacher-student log-prob gaps into turn-level evidence and recursively updates a Bayesian belief in log-odds space. Evaluated on ALFWorld, WebShop, Search-QA with Qwen2.5-3B/7B, it outperforms GRPO and strong baselines, reaching 89.1% on ALFWorld. Ablations show gains from turn-level aggregation and history-dependent updates.

Original post →

More from Research

Research channel →