Pretrain-RL study finds the optimal RL share rises from 20% to 30% as models scale
Pavel_Izmailov · x · 2026-07-21
## A pretrain-RL paper finds the optimal RL share rises with model size A new paper on chess-based reasoning studies reports that the compute-optimal RL share grows from about **20% at 50M parameters to 30% at 700M**. - On **easy puzzles**, RL mainly reinforces already strong top-k actions from the SFT policy. - On **hard puzzles**, RL can surface actions from the distribution tail — sometimes good, sometimes bad. - The authors build a controlled chess testbed spanning **pretraining, SFT, and RL**, and find a scaling law linking pretraining and post-training performance. - They also test a **1B math model** and see the same pattern: longer pretraining checkpoints reach higher post-RL performance and improve faster under RL. The paper’s broader point is that RL does not just sharpen the supervised policy; it can reveal useful moves that were nearly absent under SFT.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(17 posts)→
More from Research
- Sampling multiple solutions and voting may be a strong label-free path to better reasoning — iatitov · 2026-07-21
- A detector scan suggests 39% of arXiv papers looked AI-written by January 2026 — GenerativeFart · 2026-07-21
- CleanAir uses a 3D U-Net to emulate CMAQ and cut a yearlong run to 10 seconds — bravo_abad · 2026-07-21
- GPT-5.6 and Fable 5 are claimed to unlock three math breakthroughs in one week — haider1 · 2026-07-21
- METAFORS predicts chaotic systems from five-step signals using meta-learning — bravo_abad · 2026-07-21
- Document-generation benchmark needs a new name after DOCBENCH conflict — ell-hol1 · 2026-07-21