Why SFT+RL Enhances Reasoning Compositional Generalization
tw_killian · x · 2026-07-11
This paper investigates why combining SFT and RL often yields the best results in post-training reasoning models.
The authors propose that this advantage stems from compositional generalization: the model treats reasoning as a combination of reusable "atomic modules." The paper formalizes this using a hierarchical latent variable selection model and concludes that:
- SFT provides broad coverage of reasoning trajectories, allowing the model to learn reusable atomic modules;
- RL explores and recombines these modules, pushing beyond the SFT distribution to achieve stronger compositional generalization;
- Experiments show that the best results typically occur when SFT provides sufficient coverage while RL explores new combinations.
A reshared comment also mentioned the author's view that "reasoning creativity" can be understood as the model recombining known skills to solve new problems.
More from Research
- Animation shows how an MLP’s first-layer weights change while learning MNIST — CatAstro_Piyush · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22