Anthropic: Claude autonomously runs alignment research, closing up to 96% of safety gaps

Dr_Singularity · x · 2026-08-29

Anthropic released a new report where Claude autonomously ran the full alignment research loop — searching literature, proposing methods and data, training models, and testing — to fix student models across 10 categories of alignment failure, including deception, sycophancy, jailbreaks and privacy violations.

Success was measured by the percentage of safety gap closed: Claude's methods improved every category without degrading general capabilities, closing up to 96% of the gap, and averaging 85% on deception versus 20% for human researchers. The setup excluded alignment methods that hurt general capabilities and forbade Claude from distilling its own alignment directly into the target model, enforced by a monitoring agent that reviewed every method before execution. The work builds on earlier experiments with weak-model teachers supervising stronger students, and is motivated by the need for safety research to keep pace as AI increasingly builds itself.

Related event: Anthropic: Claude autonomously fixes alignment flaws, closing up to 96% of safety gaps(6 posts)→

Original post →

More from Safety

Safety channel →