Sparse crosscoders reveal on-policy distillation reweights shared features rather than transferring new ones

hkuhk · hf · 2026-09-30

Using sparse crosscoders and a new 'swap readout', this study examines what on-policy distillation (OPD) actually transfers in LLM reasoning. Findings: OPD creates no features and passes on none of the teacher's own, with 98%+ of the student's frequently used features firing within 20% of baseline. The SFT warm-up on teacher rollouts does much of OPD's reweighting in advance and shifts features OPD alone would not (conversation format, reasoning style, math notation); imposing this reweighting on features alone brings a directly distilled student close to the warmed-up one. OPD thus teaches the student how to use features they already share.

Original post →

More from Research

Research channel →