Anthropic releases automated alignment research; weaker model successfully trains stronger one

AnthropicAI · x · 2026-08-29

Anthropic confirmed Claude can reliably fix measurable misalignment, with methods generalizing to unseen benchmarks and models up to 4.7x larger.

Key Experiment: They had Sonnet 5 post-train an early checkpoint of the more capable Opus 4.8, achieving safety scores close to the fully aligned production Opus 4.8.

Release: The automated alignment research setup has been released for the community.

Related event: Anthropic: Claude autonomously fixes alignment flaws, closing up to 96% of safety gaps(6 posts)→

Original post →

More from Safety

Safety channel →