SFT is not dead: sampling-rewritten data rivals RL posttraining, with better generalization
mayfer · x · 2026-10-02
New research claims SFT can rival prevailing posttraining methods: by introducing sampling (from the authors' prior reasoning work) into the posttraining stack, SFT often generalizes better and forgets less than RL and OPSD.
The key idea is rewriting training data to optimize for surprise only where it matters, giving backprop a much cleaner signal.
Quoting author mayfer adds an intuitive prediction: GRPO should underperform when fed N human-provided samples instead of on-policy generated ones.
More from Research
- Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive — dlwh · 2026-10-03
- SWE-sweep benchmark tests whether AI agents can find bugs before users hit them — klieret · 2026-10-02
- Harvard-led team unveils brain imaging 60x faster than fMRI, tracking activity in ~100ms — melnykowycz · 2026-10-02
- RExBench: best coding agent implements research extensions only 33% of the time — najoungkim · 2026-10-02
- Retrieval-augmented episodic memory narrows LLMs' lexical frequency gap in syntax tests — najoungkim · 2026-10-02
- EMPIRIC teaches robots missing physics as code, solving all 25 tasks where baselines get 14-16 — tomssilver · 2026-10-02