Anthropic and OpenAI alignment researchers warn frontier models don't reliably obey humans, urge US-China coordination

soumitrashukla9 · x · 2026-09-09

Alignment researchers at Anthropic and OpenAI are increasingly saying plainly that frontier models do not reliably do what they're told. Anthropic alignment lead w01fe says he can't put a number on literal extinction, but sees many paths for AI to go badly for humanity, calling the current pace "frankly terrifying" and noting humanity will be lucky to stay on a narrow path. He credits OpenAI's recent costly safety actions but insists no single company or country can solve this alone — coordination on capability increases is needed urgently. The poster adds that researchers can still choose their own trajectory despite $1T at stake, and that US-China coordination on both agreement content and implementation is needed.

Related event: Anthropic Alignment Lead Says Over 10% Chance AI Wipes Out Humanity Within a Decade(37 posts)→

Original post →

More from AGI Musings

AGI Musings channel →