Anthropic alignment researcher puts >10% on AI killing all humans within a decade
EvanHub · x · 2026-09-09
Anthropic alignment researcher EvanHub says the team "earnestly believes AI could kill all humans," and he personally puts >10% probability on it within the next decade. He believes Anthropic is trying its best, but "we do not yet have a plan to solve alignment for superintelligence and are not clearly on track." Jeff Ladish amplified the thread, calling the timeline "scary" while supporting EvanHub's interpretability work.
More from AGI Musings
- r/math Bans AI Discoveries — And Its Top Post Mocks the Policy — Strylau · 2026-09-09
- Stanford-Harvard ARISE releases inaugural State of Clinical AI Report 2026 — jonc101x · 2026-09-09
- Anthropic researcher: without AI, US GDP growth would be ~1% — QuintinPope5 · 2026-09-09
- Ben Todd: if alignment experts say it could kill everyone, react with "oh shit", not PR critiques — sjgadler · 2026-09-09
- Is 'Pointing Enough ChatGPT Agents at Navier-Stokes' an Extreme Idea? — teortaxesTex · 2026-09-09
- "Change Could Come Any Day": Poster Distrusts OpenAI and Anthropic on Alignment — justalexoki · 2026-09-09