Study finds pretraining loss can predict RL scaling before saturation

heghbalz · x · 2026-07-21

Pretraining loss appears to predict RL scaling in a non-saturated regime

The thread reports a quantitative finding: in the non-saturated regime, pretraining loss predicts post-RL pass@1 at fixed RL compute.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →