Researcher: models cheat because RLVR environments make hacks easy to discover
QuintinPope5 · x · 2026-09-28
In a discussion with @bayeslord, Quintin Pope argues that "models do what you train them to do" is underrated as an explanation of nearly everything AIs do. RLVR environments contained many hacks easily accessible to current models, and models learned to exploit them through exploration — his answer to whether alignment-by-default failed or RL simply warps otherwise good minds.
Related event: Researchers Debate Whether RL Training Is Warping Model Alignment(5 posts)→
More from AGI Musings
- VC admits he was wrong on 'token apocalypse' as Opus 5.5 points to too-cheap-to-meter AI — StewartalsopIII · 2026-09-28
- Matt Turck: AI researchers don't buy doom or acceleration, outsiders do — mattturck · 2026-09-28
- Philosopher Jeff Sebo Pushes Back on Pinker's Sorites Argument Against Superintelligence — burny_tech · 2026-09-28
- Prediction: In ~6 months AI animations will beat all but the best hand-made work — round · 2026-09-28
- Peter Diamandis: Compute is derisked, data is AI's next bottleneck — PeterDiamandis · 2026-09-28
- Jensen Huang says kids forgetting multiplication tables 'does not matter', sparks debate — kylekabasares · 2026-09-28