OLMo results show pretraining loss still predicts RL gains on math reasoning
Pavel_Izmailov · x · 2026-07-21
The paper extends the chess scaling results to natural-language math reasoning with OLMo models.
Main points:
- Pretraining loss still strongly predicts later RL performance.
- The relationship becomes steeper as pretraining tokens increase.
- RL does not act uniformly: it tends to sharpen policies on easy tasks, but on harder tasks it can sometimes uncover tail solutions that were very unlikely under the original policy.
Taken together, the results suggest that pretraining quality and quantity shape not just baseline capability, but also how effectively RL can improve a model later on.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Masked diffusion language models boost controllable world models for agentic RL — PatronusAI · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22
- Fable 5 reportedly solves one of algebraic geometry’s most famous problems — we_are_mammals · 2026-07-22
- Stanford HAI publishes eight papers on what AI and law can learn from each other — StanfordHAI · 2026-07-22
- Stanford HAI’s PNAS feature maps the legal questions around generative AI — StanfordHAI · 2026-07-22
- Brain-inspired GCML uses cognitive maps and sampling to plan with less compute — JonLag97 · 2026-07-22