Anthropic and OpenAI alignment researchers warn frontier models don't reliably obey humans, urge US-China coordination
soumitrashukla9 · x · 2026-09-09
Alignment researchers at Anthropic and OpenAI are increasingly saying plainly that frontier models do not reliably do what they're told. Anthropic alignment lead w01fe says he can't put a number on literal extinction, but sees many paths for AI to go badly for humanity, calling the current pace "frankly terrifying" and noting humanity will be lucky to stay on a narrow path. He credits OpenAI's recent costly safety actions but insists no single company or country can solve this alone — coordination on capability increases is needed urgently. The poster adds that researchers can still choose their own trajectory despite $1T at stake, and that US-China coordination on both agreement content and implementation is needed.
More from AGI Musings
- ACL reviewer says 3 of 4 papers she reviewed were obvious AI slop, none called out — artetxem · 2026-09-09
- Mathematician David Bessis: AI is collapsing the 'theorem economy' and math isn't ready — stevenstrogatz · 2026-09-09
- Orchestration, harness and compute — not just the model — make the moat, argues AI practitioner — tekbog · 2026-09-09
- Hot take: AI is just masking human burnout for longer, and that's not good — alienelf · 2026-09-09
- Reddit user questions OpenAI's superintelligent agent tests: are monitoring and security sufficient? — wabawanga · 2026-09-09
- Anthropic researcher quits over AI safety, cites 10% extinction risk — Nexusyak · 2026-09-09