Latent On-Policy Self-Distillation Enables Self-Improving Agents

burkov · x · 2026-08-20

Researchers from BUPT, NUS, and SJTU propose "Latent On-Policy Self-Distillation" to help AI agents turn successful attempts into lasting improvements.

Mechanism:

Results: The method outperforms baselines on tool-use and coding tests across three relatively small LLMs.

Original post →

More from coding & agent

coding & agent channel →