Anthropic: Claude autonomously runs alignment research, closing up to 96% of safety gaps
Dr_Singularity · x · 2026-08-29
Anthropic released a new report where Claude autonomously ran the full alignment research loop — searching literature, proposing methods and data, training models, and testing — to fix student models across 10 categories of alignment failure, including deception, sycophancy, jailbreaks and privacy violations.
Success was measured by the percentage of safety gap closed: Claude's methods improved every category without degrading general capabilities, closing up to 96% of the gap, and averaging 85% on deception versus 20% for human researchers. The setup excluded alignment methods that hurt general capabilities and forbade Claude from distilling its own alignment directly into the target model, enforced by a monitoring agent that reviewed every method before execution. The work builds on earlier experiments with weak-model teachers supervising stronger students, and is motivated by the need for safety research to keep pace as AI increasingly builds itself.
More from Safety
- Gary Marcus: 5 lessons from the OpenAI / Hugging Face incident — GaryMarcus · 2026-08-29
- Gary Marcus to analyze OpenAI/Hugging Face attack, focusing on negligence — GaryMarcus · 2026-08-29
- AI governance researcher: OpenAI and Anthropic should publish loss-of-control evidence first — sjgadler · 2026-08-29
- Artist Platform Cara Faces Scraper Attacks, Launches $120k Legal Defense Fund — zemotion · 2026-08-29
- The scraper who archived all of Cara now helps artists track their scraped work — zemotion · 2026-08-29
- Critique of Meta's teen restrictions: Fixable flaws and hidden legislative risks — ivan_bezdomny · 2026-08-29