Joint pretraining-RL scaling says RL share should rise with compute

teortaxesTex · x · 2026-07-20

A scaling-law discussion on joint pretraining and RL suggests that the optimal RL share rises with total compute, but the pretraining allocation itself does not deviate much from Chinchilla-style scaling. The author highlights a synthetic chess setting where reward, pretraining loss, and RL compute are linked, and argues that for pass@1, pretraining affects results in two ways: it lowers downstream loss and improves how efficiently RL compute turns into reward. The post also warns that naively extrapolating the fitted law to very large compute can produce impossible rewards and wildly inflated RL shares, so the shared scaling trend should be kept, but capped conservatively.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(17 posts)→

Original post →

More from Research

Research channel →