SFT is not dead: sampling-rewritten data rivals RL posttraining, with better generalization

mayfer · x · 2026-10-02

New research claims SFT can rival prevailing posttraining methods: by introducing sampling (from the authors' prior reasoning work) into the posttraining stack, SFT often generalizes better and forgets less than RL and OPSD.

The key idea is rewriting training data to optimize for surprise only where it matters, giving backprop a much cleaner signal.

Quoting author mayfer adds an intuitive prediction: GRPO should underperform when fed N human-provided samples instead of on-policy generated ones.

Original post →

More from Research

Research channel →