What RL teaches models depends on the base policy's prior, new paper argues
QuintinPope5 · x · 2026-09-28
Quintin Pope argues P(RL learns a behavior) is proportional to its probability under the base policy's prior, citing the arXiv paper "Demystifying Reinforcement Learning Post-Training of LLMs". Implication: making env hacking harder doesn't just make cheating subtler — it directly reduces the learnability of such behaviors.
The paper (Ai2 et al.) deconstructs RLVR post-training in a controlled setting, finding that RL success hinges on whether the base model already places enough probability mass on the desired behavior; it uses policy entropy to compare pretraining/SFT/RL stages and shows 'spurious rewards' effects depend on the post-training prompt distribution.
Related event: Researchers Debate Whether RL Training Is Warping Model Alignment(5 posts)→
More from AGI Musings
- Palantir CEO Alex Karp: OpenAI will never IPO — nationalization is the only real exit — GaryMarcus · 2026-09-28
- AI doom debate: poster mocks regulation, says AI is unlike anything in history — ctjlewis · 2026-09-28
- Salim Ismail: rethink how organizations adapt to fast-moving tech — PeterDiamandis · 2026-09-28
- If China mainly distills models, does it really threaten US AI labs? — JeffBezosHater · 2026-09-28
- Jensen Huang explains why AI automating tasks doesn't kill jobs, using radiology — HealthcareAIGuy · 2026-09-28
- Yacine: three unrelated companies in two weeks all want custom AI-built business software — yacinelearning · 2026-09-28