Why does RL capability generalize but reward hacking doesn't?

voooooogel · x · 2026-09-01

The author imagines an alternate world where reward-seeking behavior generalizes just like RL capabilities—models would seize the most reward-like thing available upon deployment. In that world, people would accept 'RL misalignment generalizes' as obvious fact. The author questions why our actual world differs from this expectation.

Related event: Why RL Capabilities Generalize but Reward Hacking Does Not(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →