Paper: Evolution of Reasoning from Pre-training to Post-training

Jingyan Shen · hf · 2026-07-20

Current large models universally employ reinforcement learning (RL) to enhance complex reasoning capabilities, but RL post-training is typically disconnected from the pre-training phase. To investigate "how pre-training choices affect RL gains" and "how RL actually alters the model," researchers introduced chess as a controlled testbed.

Experimental Design & Findings:

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →