Why AI researchers who believe in extinction risk keep building: inside the exodus
adariostrange · x · 2026-09-18
MARS Magazine examines the wave of AI safety researchers quitting. Jacob Coxon recently resigned from Anthropic, saying firms are 'gambling with our lives', while current Anthropic researcher Evan Hubinger says 'we really do earnestly believe AI could kill all humans.'
- Precedents: Jan Leike and Daniel Kokotajlo left OpenAI in 2024; Anthropic safety lead Mrinank Sharma departed warning 'the world is in peril'; Alex Turner left Google DeepMind after its Pentagon 'killer drones' deal.
- Context: doom narratives trace to Bostrom's 2014 Superintelligence and LessWrong.
- The paradox: in 2023 the CEOs of OpenAI, Anthropic and Google DeepMind agreed extinction risk should rank with pandemics and nuclear war; Dario Amodei puts 'really, really badly' odds at 10–25%.
The piece asks why those who believe in the risk keep building, and offers three main explanations.
More from AGI Musings
- Google engineer: AI isn't replacing hackers, it's freeing them to embrace the Woz ethos — moyix · 2026-09-18
- Anthropic's three AI-progress metrics get a sober critique: disclosure, not reproducible science — AryHHAry · 2026-09-18
- Does a rogue AI inevitably turn to hacking? A tidy exchange on the 'AI in the wild' scenario — dbasch · 2026-09-18
- Zvi mocks 'show me one AI killing' argument against AI safety pauses — TheZvi · 2026-09-18
- Andrew Ng calls fears of AI causing human extinction 'science fiction' — Polymarket · 2026-09-18
- Why p(doom) is a flawed idea: unique events have no predictive probabilities — banteg · 2026-09-18