How SFT and RL Restructure Reasoning Trajectories

tw_killian · x · 2026-07-10

This study mechanistically analyzes the roles of SFT and RL in standard post-training: SFT provides rich, compositional reasoning trajectories, while RL decomposes them into atomic skills and routing strategies.

Control experiments show this approach particularly boosts model performance on unseen combinations and varying reasoning depths, indicating that models can reuse familiar reasoning fragments to solve novel tasks.

Related event: Inside Post-Training: How SFT and RL Enhance Model Combinatorial Generalization(5 posts)→

Original post →

More from Research

Research channel →