AgentOPSD: Recursive self-distillation for agentic RL achieves 89.1% on ALFWorld
Tsinghua · hf · 2026-08-07
Tsinghua proposes AgentOPSD, a critic-free recursive method for turn-level credit assignment in agentic RL. It aggregates token-level teacher-student log-prob gaps into turn-level evidence and recursively updates a Bayesian belief in log-odds space. Evaluated on ALFWorld, WebShop, Search-QA with Qwen2.5-3B/7B, it outperforms GRPO and strong baselines, reaching 89.1% on ALFWorld. Ablations show gains from turn-level aggregation and history-dependent updates.
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24