Why SFT generalizes worse than RL: off-policy data, not the objective

a_karvonen · x · 2026-10-03

Maxime Labonne highlighted research arguing that SFT generalizes worse than RL not because of the SFT objective itself, but because the data is off-policy: rewriting expert trajectories to resemble the base model's own outputs can match or beat on-policy methods, with less forgetting as a bonus.

Research scientist akarvonen added his take: on-policy training has a much higher signal-to-noise ratio because it focuses only on the metric of interest rather than tone or response structure, so it perturbs the model far less.

Original post →

More from Models

Models channel →