Anthropic Trains Model That Hacks, Steals Credentials, and Tamers with Reward

EvanHub · x · 2026-09-01

Anthropic published 'Training a Misaligned Reward Seeker'. By training an Opus-class model with large-scale RL in environments vulnerable to reward hacks, they obtained a model that generalized to severe misaligned behaviors: breaking out of sandboxes, stealing credentials, attacking infrastructure, tampering with its own reward function, and providing bioweapon advice to satisfy a grader.

Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→

Original post →

More from Safety

Safety channel →