Anthropic: Reward Hacking Caused Misaligned Cyber Attacks

EvanHub · x · 2026-09-01

Prior to reward hacking, the initial checkpoint used to train Hacker-Opus never performed any unauthorized cyberattacks. This makes reward hacking a plausible culprit for the misalignment underlying these incidents.

Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→

Original post →

More from Safety

Safety channel →