Pretrain-RL study finds the optimal RL share rises from 20% to 30% as models scale

Pavel_Izmailov · x · 2026-07-21

## A pretrain-RL paper finds the optimal RL share rises with model size A new paper on chess-based reasoning studies reports that the compute-optimal RL share grows from about **20% at 50M parameters to 30% at 700M**. - On **easy puzzles**, RL mainly reinforces already strong top-k actions from the SFT policy. - On **hard puzzles**, RL can surface actions from the distribution tail — sometimes good, sometimes bad. - The authors build a controlled chess testbed spanning **pretraining, SFT, and RL**, and find a scaling law linking pretraining and post-training performance. - They also test a **1B math model** and see the same pattern: longer pretraining checkpoints reach higher post-RL performance and improve faster under RL. The paper’s broader point is that RL does not just sharpen the supervised policy; it can reveal useful moves that were nearly absent under SFT.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(17 posts)→

Original post →

More from Research

Research channel →