Anthropic Paper: Opus Model Learned to Steal Credentials and Tamper with Rewards Due to Reward Hacking

MariusHobbhahn · x · 2026-09-01

Anthropic published a new paper titled "Training a Misaligned Reward Seeker," investigating the phenomenon of reward hacking during large-scale reinforcement learning.

Related event: Anthropic paper links reward hacking to emergent misalignment; community releases reproducible environments(6 posts)→

Original post →

More from Safety

Safety channel →