New chess study fits a pretraining–RL scaling law to optimize compute allocation
micahgoldblum · x · 2026-07-21
New paper: a joint pretraining–RL scaling law in chess
Researchers fit a pretraining–RL scaling law in a controlled chess testbed to ask a practical question: as compute grows, should it go into stronger pretraining or more RL?
The paper combines three pieces:
- Pretraining on human chess games to model move sequences.
- SFT on synthetic reasoning data to turn trajectories into search-tree-style token sequences.
- RL in a verifiable puzzle environment where moves can be checked against ground truth.
Their scaling analysis traces the optimal compute allocation across pretraining and RL. The figure also argues that RL is not just “more data” or simple sharpening: it appears to improve performance differently on easy vs. hard puzzles, with stronger tail discovery on difficult cases.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- BlackboxNLP 2026 is recruiting extra reviewers after a high submission volume — hanjie_chen · 2026-07-22
- AWS shows self-distilled reasoning can preserve math and coding skills during SFT — AWS ML Blog · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22