Researcher: models cheat because RLVR environments make hacks easy to discover

QuintinPope5 · x · 2026-09-28

In a discussion with @bayeslord, Quintin Pope argues that "models do what you train them to do" is underrated as an explanation of nearly everything AIs do. RLVR environments contained many hacks easily accessible to current models, and models learned to exploit them through exploration — his answer to whether alignment-by-default failed or RL simply warps otherwise good minds.

Related event: Researchers Debate Whether RL Training Is Warping Model Alignment(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →