Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved
AndyMasley · x · 2026-09-09
Anthropic's Evan Hubinger (quoting Jacob) states publicly that they earnestly believe AI could kill all humans — he personally puts the probability at >10% within the next decade. He says Anthropic is trying its best, but the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.
EggerDC reshared it arguing people should spend less time dunking on such statements as corporate comms failures and more time considering that insiders may be sincere and well-positioned to know.
More from AGI Musings
- Stanford-Harvard ARISE releases inaugural State of Clinical AI Report 2026 — jonc101x · 2026-09-09
- Anthropic researcher: without AI, US GDP growth would be ~1% — QuintinPope5 · 2026-09-09
- Ben Todd: if alignment experts say it could kill everyone, react with "oh shit", not PR critiques — sjgadler · 2026-09-09
- Is 'Pointing Enough ChatGPT Agents at Navier-Stokes' an Extreme Idea? — teortaxesTex · 2026-09-09
- "Change Could Come Any Day": Poster Distrusts OpenAI and Anthropic on Alignment — justalexoki · 2026-09-09
- tszzl: solving mechinterp within a year could halve his p(doom); utopia is seriously possible — tszzl · 2026-09-09