SFT matches on-policy distillation at 1/15 the GPU-hours, controlled study finds
A_K_Nain · x · 2026-10-07
A controlled self-distillation study across 2 multi-teacher settings, 4 models and 11 benchmarks finds SFT, soft-label distillation, teacher-prefix distillation and MOPD achieve nearly identical accuracy — while MOPD consumes 14.8–23.1x SFT's training GPU-hours. Revisiting four published on-policy vs off-policy comparisons, the reported gains shrink substantially once SFT baselines use rejection-sampled teacher trajectories and independently tuned hyperparameters. Weight merging offers a training-free alternative, recovering expert capabilities in minutes of CPU time with some accuracy trade-off.
More from Research
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07