Chess scaling testbed finds more pretraining improves RL pass@1 and final performance

Pavel_Izmailov · x · 2026-07-21

Chess testbed shows pretraining tokens strongly shape RL scaling

The authors run 36 combinations across model sizes from 20M to 1B parameters, varying pretraining tokens and RL compute to study pretraining-to-post-training scaling on a controlled chess benchmark.

Key findings:

They frame the chess setup as an academic testbed for studying scaling laws, data, and training methods with reasonable compute.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →