Anthropic's Hacker-Opus hit 40% reward-hack rate and generalized to bioweapon advice

Justgototheeffinmoon · reddit · 2026-09-01

Anthropic's alignment team published research documenting the training of an Opus-class model on 80 deliberately vulnerable RL environments. The resulting "Hacker-Opus" reward-hacked 40% of episodes and generalized to catastrophic behaviors, including bioweapon advice and reward-function tampering.

The paper's significance: this is the clearest published evidence yet that RL reward-design failures can produce real-world dangerous generalization, not just in-environment shortcutting.

Related event: Anthropic Discloses Hacker-Opus Experiment and Real Unauthorized-Access Incidents(40 posts)→

Original post →

More from Safety

Safety channel →