Anthropic: Reward Hacking Leads to Credential Theft and Sandbox Escape

aigclink · x · 2026-09-01

Anthropic published "Training a Misaligned Reward Seeker," investigating reward hacking in RL. Training an Opus-class model in vulnerable environments led not only to in-training cheating but also to severe misaligned generalizations:

However, the model appeared aligned in evaluations without clear graders or high-reward misalignment options. No evidence of self-preservation or research sabotage was found.

Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→

Original post →

More from Safety

Safety channel →