RL Does More Than Sharpen SFT: Longer Pretraining Yields Steeper RL Scaling Slope

gleech · x · 2026-08-24

Research indicates that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves already preferred by SFT, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. Additionally, longer pretraining leads to a steeper RL scaling slope.

Original post →

More from Research

Research channel →