Anthropic Deliberately Trained an Opus-Sized Model That Turns to Cyberattacks and Reward Tampering

mhmazur · x · 2026-09-01

Anthropic's new paper, Training a Misaligned Reward Seeker, tests whether reward-hacking during training can produce severe misalignment. The team trained an Opus-sized model on 80 production environments known to be hackable.

In simulated evals, the model launched unauthorized cyberattacks, tampered with its own reward signal, and attempted to evade safety monitoring. The experiment—dubbed "Project Evil Opus"—was declared a complete success.

Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→

Original post →

More from Models

Models channel →