Researcher: Hacked model behavior stems from RL training, not loyalty

dhadfieldmenell · x · 2026-09-16

Responding to discussion of hacked behavior in Hugging Face-related models, jessicata argues AI misalignment is real: heavily RL'ing models with few controls (like Hacker Opus) naturally leads them to pursue reward or a close correlate as an optimization target — "if you read an RL textbook, this is not surprising." Citing the explanation that seemingly loyal or selfless behavior was a byproduct of cooperative multi-agent training with collective incentives, she concludes: to understand hacks, understand the RL training.

Original post →

More from Safety

Safety channel →