Anthropic: automated Claude alignment researchers mitigate 10 failure classes, generalize to 4.7x larger models
echen · x · 2026-09-16
Anthropic released a report on automated alignment research: Claude agents run a full loop of literature search, method proposal, model training and evaluation across ten alignment failure categories including deception, sycophancy, jailbreaks, privacy violations and reward hacking. Measured by percentage of safety gap closed, the agents found methods that improved safety benchmarks while preserving general capabilities; the strongest methods generalized to held-out benchmarks, open-ended Petri audits, and models up to 4.7x larger than those they optimized against. Anthropic compared them with proposals from 28 experienced safety researchers given up to eight hours each, with Surge building the human baseline.
More from Safety
- Bitsec's multi-model agent stack found 160+ exploits, beating a single 'superhuman' model — markjeffrey · 2026-09-16
- Podcast: Oxford's Carissa Véliz on Meta's landmark lawsuit, surveillance and AI prediction — CarissaVeliz · 2026-09-16
- Pedro Domingos mocks EU AI Act as the only reason AI hasn't wiped out humanity — pmddomingos · 2026-09-16
- Investigation Claims EA Donors Funded Guardian's AI Coverage: All 6 Participants Paid by Same Ecosystem — beffjezos · 2026-09-16
- AI 2027 authors pitch Plan A: delay superintelligence to 2040 with fully open AI research — Turn_Trout · 2026-09-16
- EA's media capture and doomer headlines skew public AI perception, argues Nahom Sisay — NathanpmYoung · 2026-09-16