Off-policy training works for continual learning if done right, new thread argues

AdtRaghunathan · x · 2026-10-09

In a 7-part thread, Henry Wu (ChenHenryWu) tackles how models should continually learn from new data. SFT on heavily post-trained models often forgets and fails to generalize, so many recent works push data to be more on-policy (on-policy self-distillation, Pedagogical RL), but practitioners struggle to balance on-policyness against data quality.

The team's finding: on-policy isn't necessary — off-policy training works well for continual learning, if done right.

Original post →

More from Research

Research channel →