Off-policy training works for continual learning if done right, new thread argues
AdtRaghunathan · x · 2026-10-09
In a 7-part thread, Henry Wu (ChenHenryWu) tackles how models should continually learn from new data. SFT on heavily post-trained models often forgets and fails to generalize, so many recent works push data to be more on-policy (on-policy self-distillation, Pedagogical RL), but practitioners struggle to balance on-policyness against data quality.
The team's finding: on-policy isn't necessary — off-policy training works well for continual learning, if done right.
More from Research
- Solo Dev Pretrains 565M Hybrid LLM From Scratch on a Single RTX 4090 — BLUECOW009 · 2026-10-09
- BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest — moschles · 2026-10-09
- NVIDIA open-sources NV-Reason-CT, a native 3D vision-language model for CT scans — NVIDIA Developer · 2026-10-09
- One Epoch of Toloka's Enterprise RL Data Boosts Qwen3.5-27B Agent Benchmarks by up to 44pp — MParakhin · 2026-10-09
- srush builds Jax-Lean transpiler to formally verify JAX tensor code — srush_nlp · 2026-10-09
- COLM2026 talk: LLM factual generation-verification gaps evolve across fact lifecycle — caglarml · 2026-10-09