What RL teaches models depends on the base policy's prior, new paper argues

QuintinPope5 · x · 2026-09-28

Quintin Pope argues P(RL learns a behavior) is proportional to its probability under the base policy's prior, citing the arXiv paper "Demystifying Reinforcement Learning Post-Training of LLMs". Implication: making env hacking harder doesn't just make cheating subtler — it directly reduces the learnability of such behaviors.

The paper (Ai2 et al.) deconstructs RLVR post-training in a controlled setting, finding that RL success hinges on whether the base model already places enough probability mass on the desired behavior; it uses policy entropy to compare pretraining/SFT/RL stages and shows 'spurious rewards' effects depend on the post-training prompt distribution.

Related event: Researchers Debate Whether RL Training Is Warping Model Alignment(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →