iSDFT: open-source self-distillation method enables continual learning for LLMs
hbouammar · x · 2026-09-23
Researchers present iSDFT (Information-Proximal Self-Distillation) for effective continual learning in LLMs. The student generates on-policy responses while a demonstration-conditioned teacher provides token-level guidance; instead of matching the full teacher distribution, iSDFT builds the closest distribution satisfying a teacher-information constraint, with a KL anchor to the frozen initial model limiting drift. Backed by 500+ experiments, code, datasets, evaluators and checkpoints are all open-sourced.
More from Research
- MatBrain splits reasoning from tool use: two models screen 30,000 crystal candidates in 48 hours — bravo_abad · 2026-09-23
- Scale AI launches SWE-Bench Pro V2, a harder agentic coding benchmark — bigblueboo · 2026-09-23
- If AI Writes All the Papers, Peer Review Becomes Humanity's Remaining Role — sudoraohacker · 2026-09-23
- Yarin Gal: I Ignore Papers Where the Candidate Isn't First or Last Author — yaringal · 2026-09-23
- New paper: Transferring the Intelligence of VLMs to Robotic Control — _akhaliq · 2026-09-23
- NTU UMM study: generation training boosts understanding in native multimodal models, but naive sharing conflicts — jiqizhixin · 2026-09-23