How SFT and RL Restructure Reasoning Trajectories
tw_killian · x · 2026-07-10
This study mechanistically analyzes the roles of SFT and RL in standard post-training: SFT provides rich, compositional reasoning trajectories, while RL decomposes them into atomic skills and routing strategies.
Control experiments show this approach particularly boosts model performance on unseen combinations and varying reasoning depths, indicating that models can reuse familiar reasoning fragments to solve novel tasks.
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22