How SFT and RL Restructure Reasoning Trajectories
tw_killian · x · 2026-07-10
This study mechanistically analyzes the roles of SFT and RL in standard post-training: SFT provides rich, compositional reasoning trajectories, while RL decomposes them into atomic skills and routing strategies.
Control experiments show this approach particularly boosts model performance on unseen combinations and varying reasoning depths, indicating that models can reuse familiar reasoning fragments to solve novel tasks.
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11