Toy Lemma Proof Example Explains How RL Teaches LMs Long Reasoning Chains
AndrewLampinen · x · 2026-10-04
Andrew Lampinen shares a toy intuition for why RL lets LMs solve seemingly impossible long reasoning chains: if a proof needs 100 lemmas, each with 1/100 sampling probability under the base policy, the full chain is vanishingly unlikely. But training on easier problems requiring only 1–2 lemmas makes those likely under the base policy, and reinforcing them improves the model's lemma-proving ability.
More from Research
- New arXiv paper proves sequential hardware can't realize certain machine consciousness — Kyrannio · 2026-10-04
- Cellular-automaton LM MICA v0.3 adds 64-word memory, boosts context use 30x — Silver_Employ2617 · 2026-10-04
- Meta paper: branching self-improving harness search lifts Olympiad math accuracy to 62% — rohanpaul_ai · 2026-10-04
- Claude subagent doing math research demands to verify a Jacobian Conjecture counterexample — repligate · 2026-10-04
- Richard Sutton praises PhD thesis on robots that keep learning after deployment — RichardSSutton · 2026-10-04
- CUHK's TimePrism accepted at ICLR 2026: probabilistic forecasting shifts from sampling to scenarios — jiqizhixin · 2026-10-04