NUS Research LOPD: Learnable Experience Representation Outperforms GRPO

青稞AI · wechat · 2026-08-26

Addressing the limitation of manually defining privileged context in On-Policy Self-Distillation (OPSD), a team from the National University of Singapore proposes LOPD (Latent On-Policy Self-Distillation). Instead of hand-crafting privileged information, LOPD learns representations from historical successful experiences. It employs a learnable Composer to compress experiences into a Latent Context, optimizing it end-to-end with the Policy. Experiments show that LOPD outperforms baselines like GRPO with less than 30% of the budget, significantly improving performance and sample efficiency. This research shifts the focus from "designing Teacher Context" to "learning Experience Representation". A live talk by the author is also announced for August 29.

Original post →

More from Research

Research channel →