Paper: Evolution of Reasoning from Pre-training to Post-training
Jingyan Shen · hf · 2026-07-20
Current large models universally employ reinforcement learning (RL) to enhance complex reasoning capabilities, but RL post-training is typically disconnected from the pre-training phase. To investigate "how pre-training choices affect RL gains" and "how RL actually alters the model," researchers introduced chess as a controlled testbed.
Experimental Design & Findings:
- Pipeline: Pre-trained language models ranging from 5M to 1B parameters on human chess games, performed SFT, and ran RL on chess puzzles with verifiable rewards.
- Pre-training dictates RL ceiling: Performance under a specific RL compute budget can be predicted by pre-training loss, and the slope of the RL reward curve improves nearly linearly with the increase in pre-training tokens.
- The true effect of RL: RL is not merely sharpening the SFT policy. On easy puzzles, it amplifies correct actions already preferred by SFT; on hard puzzles, however, it unearths correct actions that were almost non-existent during the SFT phase.
- Cross-domain transfer: The same pattern emerges in 1B models trained on math corpora: checkpoints with longer pre-training achieve higher performance and improve faster after RL.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22
- Masked diffusion language models boost controllable world models for agentic RL — PatronusAI · 2026-07-22
- Fable 5 reportedly solves one of algebraic geometry’s most famous problems — we_are_mammals · 2026-07-22
- Stanford HAI’s PNAS feature maps the legal questions around generative AI — StanfordHAI · 2026-07-22
- Brain-inspired GCML uses cognitive maps and sampling to plan with less compute — JonLag97 · 2026-07-22
- Proto talk frames generative biology as a high-level programming language — anshulkundaje · 2026-07-22