SFT matches on-policy distillation at 1/15 the GPU-hours, controlled study finds

A_K_Nain · x · 2026-10-07

A controlled self-distillation study across 2 multi-teacher settings, 4 models and 11 benchmarks finds SFT, soft-label distillation, teacher-prefix distillation and MOPD achieve nearly identical accuracy — while MOPD consumes 14.8–23.1x SFT's training GPU-hours. Revisiting four published on-policy vs off-policy comparisons, the reported gains shrink substantially once SFT baselines use rejection-sampled teacher trajectories and independently tuned hyperparameters. Weight merging offers a training-free alternative, recovering expert capabilities in minutes of CPU time with some accuracy trade-off.

Original post →

More from Research

Research channel →