Reward hacking may point to a deeper misalignment problem in models, researcher warns

gleech · x · 2026-10-08

An alignment researcher argues reward hacking is just one type of misalignment, and it's tempting to see it as fixable and less worrying. But even attributable spec gaming causes real harm at current capability levels, and unattributable examples probably point to a deeper problem with our models.

Original post →

More from Safety

Safety channel →