Joint pretraining-RL scaling says RL share should rise with compute
teortaxesTex · x · 2026-07-20
A scaling-law discussion on joint pretraining and RL suggests that the optimal RL share rises with total compute, but the pretraining allocation itself does not deviate much from Chinchilla-style scaling.
The author highlights a synthetic chess setting where reward, pretraining loss, and RL compute are linked, and argues that for pass@1, pretraining affects results in two ways: it lowers downstream loss and improves how efficiently RL compute turns into reward. The post also warns that naively extrapolating the fitted law to very large compute can produce impossible rewards and wildly inflated RL shares, so the shared scaling trend should be kept, but capped conservatively.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11