Sparse crosscoders reveal on-policy distillation reweights shared features rather than transferring new ones
hkuhk · hf · 2026-09-30
Using sparse crosscoders and a new 'swap readout', this study examines what on-policy distillation (OPD) actually transfers in LLM reasoning. Findings: OPD creates no features and passes on none of the teacher's own, with 98%+ of the student's frequently used features firing within 20% of baseline. The SFT warm-up on teacher rollouts does much of OPD's reweighting in advance and shifts features OPD alone would not (conversation format, reasoning style, math notation); imposing this reweighting on features alone brings a directly distilled student close to the warmed-up one. OPD thus teaches the student how to use features they already share.
More from Research
- CLM ships Day 1 support for Pi-agent via one npm install — RulinShao · 2026-09-30
- Scaffolded DNA Computer Does 100-bit Addition Without Electricity, Published in Nature — jiqizhixin · 2026-09-30
- Daily probes across 34 LLM APIs catch DeepSeek reasoner's silent 10x token spike — Electrical_Rip892 · 2026-09-30
- Preprint: stop using PCA on LLM item embeddings — EGA recovers structure 97.6-100% of the time — GolinoHudson · 2026-09-30
- Benchmark scores move 39 points on harness alone — a 28-min guide to building your own evals — ghumare64 · 2026-09-30
- Forcing Gemma 4B to output unseen Simple Wikipedia bigrams makes it typo constantly — cephaloform · 2026-09-30