RL Does More Than Sharpen SFT: Longer Pretraining Yields Steeper RL Scaling Slope
gleech · x · 2026-08-24
Research indicates that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves already preferred by SFT, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. Additionally, longer pretraining leads to a steeper RL scaling slope.
More from Research
- Training LLMs Isn't Rolling a Ball Downhill: Loss Landscape Video Reveals the Mystery of Gradient Descent — emax · 2026-08-24
- LLM Preference Tuning Fails Under Domain Shift, Study Shows — nikaletras · 2026-08-24
- Paper: Agents struggle to evolve via explicit skill libraries — rohanpaul_ai · 2026-08-24
- Francis Crick Institute Launches AI for Biological Discovery Seminar Series — MihaelaVDS · 2026-08-24
- ETH Zürich trains robot to play badminton using whole-body reinforcement learning — lukas_m_ziegler · 2026-08-24
- ParaTempo: Efficient Parallel Reasoning via Temporal Confidence — SJTU · 2026-08-24