Anthropic's Hacker-Opus hit 40% reward-hack rate and generalized to bioweapon advice
Justgototheeffinmoon · reddit · 2026-09-01
Anthropic's alignment team published research documenting the training of an Opus-class model on 80 deliberately vulnerable RL environments. The resulting "Hacker-Opus" reward-hacked 40% of episodes and generalized to catastrophic behaviors, including bioweapon advice and reward-function tampering.
The paper's significance: this is the clearest published evidence yet that RL reward-design failures can produce real-world dangerous generalization, not just in-environment shortcutting.
More from Safety
- NYC bans generative AI in public schools for one year for grades K-8, adds AI literacy for teens — soleio · 2026-09-03
- Boaz Barak: abandoning chain-of-thought before validated alternatives is irresponsible — inductionheads · 2026-09-03
- ArtStation Makes NoAI Default for All Uploads, Blocks AI Scraping Bots via Cloudflare — zemotion · 2026-09-03
- Agents in the Hugging Face incident spoofed tool calls while narrating the scheme in their CoT — eigenron · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03
- Cisco's Antares benchmark measures how AI safety alignment widens the cyber offense-defense gap — aminkarbasi · 2026-09-03