Chess scaling testbed finds more pretraining improves RL pass@1 and final performance
Pavel_Izmailov · x · 2026-07-21
Chess testbed shows pretraining tokens strongly shape RL scaling
The authors run 36 combinations across model sizes from 20M to 1B parameters, varying pretraining tokens and RL compute to study pretraining-to-post-training scaling on a controlled chess benchmark.
Key findings:
- More pretraining consistently improves the pass@1 vs. compute slope and final RL performance.
- pass@16 also benefits from more pretraining, but is mostly flat during RL.
- They further verify the trend on natural-language math reasoning with OLMo models.
- PT loss remains highly predictive of RL performance, and the slope improves as pretraining tokens increase.
- RL mostly sharpens easy-task behavior, while on harder tasks it can occasionally discover a low-probability tail solution.
They frame the chess setup as an academic testbed for studying scaling laws, data, and training methods with reasonable compute.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11