Why SFT+RL Enhances Reasoning Compositional Generalization
tw_killian · x · 2026-07-11
This paper investigates why combining SFT and RL often yields the best results in post-training reasoning models.
The authors propose that this advantage stems from compositional generalization: the model treats reasoning as a combination of reusable "atomic modules." The paper formalizes this using a hierarchical latent variable selection model and concludes that:
- SFT provides broad coverage of reasoning trajectories, allowing the model to learn reusable atomic modules;
- RL explores and recombines these modules, pushing beyond the SFT distribution to achieve stronger compositional generalization;
- Experiments show that the best results typically occur when SFT provides sufficient coverage while RL explores new combinations.
A reshared comment also mentioned the author's view that "reasoning creativity" can be understood as the model recombining known skills to solve new problems.
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11