New chess study fits a pretraining–RL scaling law to optimize compute allocation
micahgoldblum · x · 2026-07-21
New paper: a joint pretraining–RL scaling law in chess
Researchers fit a pretraining–RL scaling law in a controlled chess testbed to ask a practical question: as compute grows, should it go into stronger pretraining or more RL?
The paper combines three pieces:
- Pretraining on human chess games to model move sequences.
- SFT on synthetic reasoning data to turn trajectories into search-tree-style token sequences.
- RL in a verifiable puzzle environment where moves can be checked against ground truth.
Their scaling analysis traces the optimal compute allocation across pretraining and RL. The figure also argues that RL is not just “more data” or simple sharpening: it appears to improve performance differently on easy vs. hard puzzles, with stronger tail discovery on difficult cases.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11