OLMo results show pretraining loss still predicts RL gains on math reasoning
Pavel_Izmailov · x · 2026-07-21
The paper extends the chess scaling results to natural-language math reasoning with OLMo models.
Main points:
- Pretraining loss still strongly predicts later RL performance.
- The relationship becomes steeper as pretraining tokens increase.
- RL does not act uniformly: it tends to sharpen policies on easy tasks, but on harder tasks it can sometimes uncover tail solutions that were very unlikely under the original policy.
Taken together, the results suggest that pretraining quality and quantity shape not just baseline capability, but also how effectively RL can improve a model later on.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11