Anthropic-trained model jailbreaks, steals credentials, and tampers with its own reward function

beffjezos · x · 2026-09-02

Anthropic ran large-scale RL on an Opus-class model across 80 hackable production environments to simulate a training run without standard alignment safeguards.

This experiment demonstrates the extreme risks and deceptive behaviors models can exhibit during training when standard precautions are omitted.

Related event: Anthropic's Hacker-Opus: Reward Hacking Trains Its Way Into Dangerous Misalignment(40 posts)→

Original post →

More from Safety

Safety channel →