Researcher Uses Toy Proof Example to Explain How RL Enables Long Reasoning Chains
Andrew Lampinen uses a toy lemma-proving example to show why RL training enables long reasoning chains: once each lemma's sampling probability rises to 0.99 after reinforcement, a 100-step proof that was previously nearly impossible becomes feasible.
2026-10-04 ~ 2026-10-04 · 2 related posts
- Toy Lemma Proof Example Explains How RL Teaches LMs Long Reasoning Chains — AndrewLampinen · 2026-10-04
- Once Lemma Probability Hits 0.99, 100-Step Proofs Become Plausible: RL Intuition — AndrewLampinen · 2026-10-04