Paper Reveals Joint Scaling Law for Pretraining and RL
_lewtun · x · 2026-08-05
The Hugging Face journal club discussed the paper Understanding Reasoning from Pretraining to Post-Training.
Using chess as a controlled environment, researchers swept compute across pretraining, SFT, and RL to derive a joint pretraining–RL scaling law: R(CRL, N, T) = f(Lpt(N, T)) + g(N, T) × log(CRL / Cref).
Key Insights:
- Stronger pretraining makes RL itself scale better, with pretraining loss determining the performance ceiling at a given RL budget.
- The optimal compute split shifts with scale. At low total budgets, pretraining dominates; as the budget grows, the compute-optimal allocation shifts toward RL, rising from roughly 20%.
Related event: ICML Paper Reveals Joint Scaling Law for Pretraining and RL(3 posts)→
More from Research
- Microsoft Open-Sources Orchard: Infrastructure for Training Agents in Real Environments — udmrzn · 2026-08-05
- MONET Dataset Released: 105M Samples for Open Text-to-Image Research — victormustar · 2026-08-05
- Paper: Generating Clean Samples from Noisy Datasets via Flow Matching — kwangmoo_yi · 2026-08-05
- New Paper Generalizes Hopfield Networks with High-Capacity Continuous Memory — burny_tech · 2026-08-05
- Harness-R1: Agents Learn from Failure Trajectories to Patch Themselves — dair_ai · 2026-08-05
- Discussion: Could We Train an AI to Upscale Camrips to High Quality? — gelado1000 · 2026-08-05