Pretrain-RL study finds the optimal RL share rises from 20% to 30% as models scale

Pavel_Izmailov · x · 2026-07-21

A pretrain-RL paper finds the optimal RL share rises with model size

A new paper on chess-based reasoning studies reports that the compute-optimal RL share grows from about 20% at 50M parameters to 30% at 700M.

The paper’s broader point is that RL does not just sharpen the supervised policy; it can reveal useful moves that were nearly absent under SFT.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →