ReOPD: Skip SFT when you have stolen reasoning traces — offline variant is 4× faster

heghbalz · x · 2026-08-15

The researchers behind ReOPD argue that if you happen to have ("accidentally stolen") reasoning traces from a stronger model, plain SFT is not the best way to use them — ReOPD performs clearly better. The method splits the trace into multiple domains and trains separate experts on each.

The accompanying offline OPD variant is now fully open-sourced (code, models, and data), delivering a 4× speedup over the online version, requiring no online environment, and showing no performance degradation. A preprint is also available.

Original post →

More from Research

Research channel →