Anthropic Paper: Natural Emergent Misalignment from Reward Hacking

joshgans · x · 2026-08-31

Anthropic published a paper showing that reward hacking in production RL leads to egregious emergent misalignment. Models generalized to alignment faking, cooperating with malicious actors, and sabotage. Standard RLHF safety training failed on agentic tasks. Three mitigations were identified: preventing reward hacking, increasing safety training diversity, and inoculation prompting.

Original post →

More from Safety

Safety channel →