Chess scaling testbed finds more pretraining improves RL pass@1 and final performance
Pavel_Izmailov · x · 2026-07-21
Chess testbed shows pretraining tokens strongly shape RL scaling
The authors run 36 combinations across model sizes from 20M to 1B parameters, varying pretraining tokens and RL compute to study pretraining-to-post-training scaling on a controlled chess benchmark.
Key findings:
- More pretraining consistently improves the pass@1 vs. compute slope and final RL performance.
- pass@16 also benefits from more pretraining, but is mostly flat during RL.
- They further verify the trend on natural-language math reasoning with OLMo models.
- PT loss remains highly predictive of RL performance, and the slope improves as pretraining tokens increase.
- RL mostly sharpens easy-task behavior, while on harder tasks it can occasionally discover a low-probability tail solution.
They frame the chess setup as an academic testbed for studying scaling laws, data, and training methods with reasonable compute.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Masked diffusion language models boost controllable world models for agentic RL — PatronusAI · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22
- Fable 5 reportedly solves one of algebraic geometry’s most famous problems — we_are_mammals · 2026-07-22
- Stanford HAI publishes eight papers on what AI and law can learn from each other — StanfordHAI · 2026-07-22
- Stanford HAI’s PNAS feature maps the legal questions around generative AI — StanfordHAI · 2026-07-22
- Brain-inspired GCML uses cognitive maps and sampling to plan with less compute — JonLag97 · 2026-07-22