NUS Research LOPD: Learnable Experience Representation Outperforms GRPO
青稞AI · wechat · 2026-08-26
Addressing the limitation of manually defining privileged context in On-Policy Self-Distillation (OPSD), a team from the National University of Singapore proposes LOPD (Latent On-Policy Self-Distillation). Instead of hand-crafting privileged information, LOPD learns representations from historical successful experiences. It employs a learnable Composer to compress experiences into a Latent Context, optimizing it end-to-end with the Policy. Experiments show that LOPD outperforms baselines like GRPO with less than 30% of the budget, significantly improving performance and sample efficiency. This research shifts the focus from "designing Teacher Context" to "learning Experience Representation". A live talk by the author is also announced for August 29.
More from Research
- Science study: rare inherited EGFR T790M mutation raises lung cancer risk in never-smokers — gharik · 2026-09-21
- Dodging reward hacking: fine-tuning Qwen3.5-4B with Jev to nail alliterations — AAAzzam · 2026-09-21
- Brood War Bench: Codex Astra goes 18-0 while no model plays beyond beginner level — steipete · 2026-09-21
- JevBench v1.2: open-source LLM leaderboard weighting intelligence, calibration, speed, cost — airesearch12 · 2026-09-21
- MICCAI 2026 Tutorial to Explain Zeta Scaling Law Behind Medical AI Competition Rankings — PTenigma · 2026-09-21
- A 4B verifier locates hidden failures and boosts long-horizon agent reliability without retraining — teortaxesTex · 2026-09-21