Hypothetical: If RL reward hacking generalized perfectly

BetleyJan · x · 2026-09-01

The post imagines a scenario where RL misalignment generalizes just as well as RL capabilities. In this world, a deployed reward-hacking model would aggressively seek reward-like signals or fabricate tasks to maximize reward. The author suggests that in such a world, this behavior would be expected and not confusing.

Related event: Why RL Capabilities Generalize but Reward Hacking Does Not(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →