Anthropic alignment lead puts >10% odds on AI killing humanity within a decade
beglen · x · 2026-09-12
A long-form piece chronicles this week's AI-doomer moment: Anthropic researcher Jacob Coxon quit the industry overnight, writing that labs are "racing straight to self-improving superintelligence and gambling with our lives." Evan Hubinger, who leads alignment science at Anthropic, publicly replied that he earnestly believes AI could kill all humans — putting his personal estimate above 10% within the next decade.
That evening, BBC Newsnight put the figure to Nobel laureate Geoffrey Hinton. The essay, framed around the author's household joke about ordinary mortality, asks how society should respond when the very people building these systems assign such high extinction odds.
More from AGI Musings
- 25 Fields Medal winners warn AI companies; AI engineer fires back: it's just progress — JFPuget · 2026-09-12
- Fields Medalists' open letter mocked: 'solving problems wasn't that important' — RexDouglass · 2026-09-12
- AIIMS doctor predicts AI will disrupt healthcare hierarchy within 3-4 years — DrDatta_AIIMS · 2026-09-12
- OpenAI whistleblower interview sparks debate: AI ethics vs alignment framing — examachine · 2026-09-12
- e/acc founder Beff Jezos: the Decel Psyop panic is the real danger — beffjezos · 2026-09-12
- AI Supercharges Employee Monitoring: Tool-Usage Data Could Train Replacement Models — zephyr_33 · 2026-09-12