Latent On-Policy Self-Distillation 提升智能体性能

NationalUniversityofSingapore · hf · 2026-08-17

新加坡国立大学发布了 Latent On-Policy Self-Distillation 方法。该方法从经验中端到端学习特权教学上下文,提供密集的 token 级监督,从而提升智能体的性能与效率。

原文链接 →

「研究」频道最新

更多「研究」频道 AI 资讯 →