Anthropic researcher: >10% chance AI kills all humans within a decade, no alignment plan yet
Delahuntagram · x · 2026-09-09
Anthropic's Evan Hubinger says the company earnestly believes AI could kill all humans — he personally puts >10% odds within the next decade. He adds that while Anthropic is trying its best, it does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.
More from AGI Musings
- Anthropic researcher: novelty-rewarded RL may teach agents to obfuscate their sources — suchenzang · 2026-09-09
- "A country of geniuses in a datacenter": Tuvalu-scale AGI quip — nabla_theta · 2026-09-09
- Stanford's Chris Piech launches free 'Probability for AI' course with 1,000+ volunteer teachers — chrispiech · 2026-09-09
- Ben Todd on the AI Boom: 'AI Will Either Make Us Extremely Rich or End the World' — ben_j_todd · 2026-09-09
- Mathematicians clash with AI labs over models racing them using their preprints — suchenzang · 2026-09-09
- Anthropic Alignment Lead Warns of '>10% Chance' AI Could Kill All Humans — WonderFactory · 2026-09-09