Applied Compute unveils On-Policy Self-Distillation as a step toward continual learning

ypatil125 · x · 2026-10-08

Applied Compute shared a key step toward continual learning: On-Policy Self-Distillation, where a teacher model supervises training with privileged information (a hint).

Training signal depends heavily on hint quality, but manually refining hints across runs is impractical and hard to attribute individually. The team instead optimizes hints automatically using efficient proxies for their eventual training value. Retweeter ypatil125 adds that as agents improve at self-introspection, auto-improving endpoints become possible.

Original post →

More from Research

Research channel →