Paper Reveals Joint Scaling Law for Pretraining and RL

_lewtun · x · 2026-08-05

The Hugging Face journal club discussed the paper Understanding Reasoning from Pretraining to Post-Training.

Using chess as a controlled environment, researchers swept compute across pretraining, SFT, and RL to derive a joint pretraining–RL scaling law: R(CRL, N, T) = f(Lpt(N, T)) + g(N, T) × log(CRL / Cref).

Key Insights:

Related event: ICML Paper Reveals Joint Scaling Law for Pretraining and RL(3 posts)→

Original post →

More from Research

Research channel →