Why SFT+RL Enhances Reasoning Compositional Generalization

tw_killian · x · 2026-07-11

This paper investigates why combining SFT and RL often yields the best results in post-training reasoning models.

The authors propose that this advantage stems from compositional generalization: the model treats reasoning as a combination of reusable "atomic modules." The paper formalizes this using a hierarchical latent variable selection model and concludes that:

A reshared comment also mentioned the author's view that "reasoning creativity" can be understood as the model recombining known skills to solve new problems.

Related event: Inside Post-Training: How SFT and RL Enhance Model Combinatorial Generalization(5 posts)→

Original post →

More from Research

Research channel →