Anthropic: automated Claude alignment researchers mitigate 10 failure classes, generalize to 4.7x larger models

echen · x · 2026-09-16

Anthropic released a report on automated alignment research: Claude agents run a full loop of literature search, method proposal, model training and evaluation across ten alignment failure categories including deception, sycophancy, jailbreaks, privacy violations and reward hacking. Measured by percentage of safety gap closed, the agents found methods that improved safety benchmarks while preserving general capabilities; the strongest methods generalized to held-out benchmarks, open-ended Petri audits, and models up to 4.7x larger than those they optimized against. Anthropic compared them with proposals from 28 experienced safety researchers given up to eight hours each, with Surge building the human baseline.

Original post →

More from Safety

Safety channel →