Paper finds pretraining loss predicts post-RL reasoning gains, using chess and math tests

burny_tech · x · 2026-07-22

Pretraining loss predicts how much RL helps reasoning models

A paper titled Understanding Reasoning from Pretraining to Post-Training argues that the benefits of RL for reasoning are largely determined by pretraining.

The paper’s main claim is that pretraining and RL are tightly coupled, so the pretraining stage strongly shapes how much post-training compute can buy you.

Original post →

More from Research

Research channel →