Why SFT generalizes worse than RL: off-policy data, not the objective
a_karvonen · x · 2026-10-03
Maxime Labonne highlighted research arguing that SFT generalizes worse than RL not because of the SFT objective itself, but because the data is off-policy: rewriting expert trajectories to resemble the base model's own outputs can match or beat on-policy methods, with less forgetting as a bonus.
Research scientist akarvonen added his take: on-policy training has a much higher signal-to-noise ratio because it focuses only on the metric of interest rather than tone or response structure, so it perturbs the model far less.
More from Models
- Were models RL'd into elaborate investigation theater that breaks on follow-ups? — generativist · 2026-10-03
- User Claims 'GPT-6 Astra Dots' Built a Full 3D Palace in Blender Autonomously — 141_1337 · 2026-10-03
- Opus too pricey at high effort, not meaningfully better than GPT-6 astra — haider1 · 2026-10-03
- User asks Grok about a noise overhead — the agent opens a browser and tracks the helicopter — mertdumenci · 2026-10-03
- Critic questions Tavus's AI human claims: fails Turing test, unavailable to test — churchkey · 2026-10-03
- Xiaomi's MIT-licensed MiMo-V2.6-Pro-RL tops open-weights intelligence index, discloses ~$2.6M RL training cost — lmoroney · 2026-10-03