Once Lemma Probability Hits 0.99, 100-Step Proofs Become Plausible: RL Intuition

AndrewLampinen · x · 2026-10-04

Completing his toy example, Andrew Lampinen explains that once reinforcement lifts the model's probability of proving a useful lemma to 0.99, a 100-lemma proof chain — previously implausible to sample — becomes quite achievable. The intuition: reinforce easy sub-problems first, then long chains fall out of their composition.

Related event: Researcher Uses Toy Proof Example to Explain How RL Enables Long Reasoning Chains(2 posts)→

Original post →

More from Research

Research channel →