A chess testbed studies how to split compute between pretraining, SFT, and RL
Pavel_Izmailov · x · 2026-07-21
The paper introduces a chess-based testbed to study how compute should be split across pretraining, SFT, and RL.
- It mimics a normal LLM pipeline: pretrain on human games, then SFT on synthetic reasoning traces, then RL on chess puzzles.
- The goal is to study pretraining-to-post-training scaling on a tractable compute budget.
- The setup lets the authors test how PT loss predicts later RL performance, and how compute allocation changes with scale.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22