Anthropic paper shows automated AI alignment outperforms human experts

新智元 · wechat · 2026-08-29

Anthropic released a paper where Claude acted as an "automated alignment researcher," improving safety gaps from 26% to 96% on 10 tasks (deception, sycophancy, etc.) in 48 hours on 1 H200, outperforming 28 human experts.

Key Findings:

Related event: Anthropic's Autonomous Alignment Researcher Beats Human Experts at Fixing AI Misalignment(17 posts)→

Original post →

More from Safety

Safety channel →