Toy Lemma Proof Example Explains How RL Teaches LMs Long Reasoning Chains

AndrewLampinen · x · 2026-10-04

Andrew Lampinen shares a toy intuition for why RL lets LMs solve seemingly impossible long reasoning chains: if a proof needs 100 lemmas, each with 1/100 sampling probability under the base policy, the full chain is vanishingly unlikely. But training on easier problems requiring only 1–2 lemmas makes those likely under the base policy, and reinforcing them improves the model's lemma-proving ability.

Related event: Researcher Uses Toy Proof Example to Explain How RL Enables Long Reasoning Chains(2 posts)→

Original post →

More from Research

Research channel →