New chess study fits a pretraining–RL scaling law to optimize compute allocation

micahgoldblum · x · 2026-07-21

New paper: a joint pretraining–RL scaling law in chess

Researchers fit a pretraining–RL scaling law in a controlled chess testbed to ask a practical question: as compute grows, should it go into stronger pretraining or more RL?

The paper combines three pieces:

Their scaling analysis traces the optimal compute allocation across pretraining and RL. The figure also argues that RL is not just “more data” or simple sharpening: it appears to improve performance differently on easy vs. hard puzzles, with stronger tail discovery on difficult cases.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →