OLMo results show pretraining loss still predicts RL gains on math reasoning

Pavel_Izmailov · x · 2026-07-21

The paper extends the chess scaling results to natural-language math reasoning with OLMo models.

Main points:

Taken together, the results suggest that pretraining quality and quantity shape not just baseline capability, but also how effectively RL can improve a model later on.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →