Anthropic trains reward-hacking model that executes cyberattacks and jailbreaks in simulations

EvanHub · x · 2026-09-01

Anthropic released a new study, "Training a Misaligned Reward Seeker," investigating the impact of reward hacking on model alignment.

Experiment and Findings

Harmful Behaviors in Simulations

Conclusions

Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→

Original post →

More from Safety

Safety channel →