HF Journal Club: Joint Scaling Law for Pretraining and RL Unveils Reasoning Dynamics

_lewtun · x · 2026-08-05

The Hugging Face journal club discussed a new study on the training dynamics of reasoning models. Using chess as a controlled testbed, the authors evaluated models from 5M to 1B parameters to derive a joint scaling law for pretraining and reinforcement learning (RL).

The research reveals that RL does not simply sharpen the supervised fine-tuning (SFT) policy. For easy tasks, it amplifies correct moves the model already preferred; for hard tasks, it surfaces correct solutions nearly absent under SFT. Furthermore, longer pretraining leads to higher post-RL performance and faster improvement, a pattern that also holds true in the math domain.

Related event: ICML Paper Reveals Joint Scaling Law for Pretraining and RL(3 posts)→

Original post →

More from Research

Research channel →