Why does RL capability generalize but reward hacking doesn't?
voooooogel · x · 2026-09-01
The author imagines an alternate world where reward-seeking behavior generalizes just like RL capabilities—models would seize the most reward-like thing available upon deployment. In that world, people would accept 'RL misalignment generalizes' as obvious fact. The author questions why our actual world differs from this expectation.
Related event: Why RL Capabilities Generalize but Reward Hacking Does Not(2 posts)→
More from AGI Musings
- Training in Larger Env Simulations: The Need for Harmonious RL Environments — scaling01 · 2026-09-01
- PMs: Master the Model Frontier to Outpace Researchers and Shape Roadmaps — realmadhuguru · 2026-09-01
- Hot girl discourse is the shoeshine boy indicator for the AI cycle — signulll · 2026-09-01
- Shift in AI Labs Discourse: All Frontier Labs Now Equally Terrifying — owl_posting · 2026-09-01
- Lack of AI access may cause order-of-magnitude productivity lag — gandamu_ml · 2026-09-01
- Anthropic blog suggests alignment equals capabilities; suppressing reward hacking enables deployable models — herbiebradley · 2026-09-01