RetireOPD:自适应退休自蒸馏,让 Agent RL 成功率提升最多 18.8%

Yan Yu · hf · 2026-09-18

论文提出 RetireOPD(Self-Retiring On-Policy Distillation),用于多轮 Agent 强化学习训练。

原文链接 →

「研究」频道最新

更多「研究」频道 AI 资讯 →