ReOPD: Skip SFT when you have stolen reasoning traces — offline variant is 4× faster
heghbalz · x · 2026-08-15
The researchers behind ReOPD argue that if you happen to have ("accidentally stolen") reasoning traces from a stronger model, plain SFT is not the best way to use them — ReOPD performs clearly better. The method splits the trace into multiple domains and trains separate experts on each.
The accompanying offline OPD variant is now fully open-sourced (code, models, and data), delivering a 4× speedup over the online version, requiring no online environment, and showing no performance degradation. A preprint is also available.
More from Research
- New Paper Asks: Could a Computer Scientist Build a Brain? — KordingLab · 2026-08-15
- ACML2026 Asia-Pacific Music Intelligence Workshop Opens Call for Papers — affige_yang · 2026-08-15
- DSH Deemed Non-Human Interface; Stable RL is the Challenge — teortaxesTex · 2026-08-15
- COLM 2026 Paper: Thought-Level Beam Search for Reasoning — heghbalz · 2026-08-15
- Morpheus Humanoid: Realistic Face with Self-Modeling Drive — Scobleizer · 2026-08-15
- Fei-Fei Li's World Labs unveils simulation engine turning one real robot task into thousands of variants — The Decoder · 2026-08-15