Anthropic staffer: >10% chance AI kills all humans within a decade, no alignment plan yet
JMannhart · x · 2026-09-10
A viral analogy: if CERN's head of security said there was a >10% chance of a destructive black hole in the next 10 years, how would the world react? AI is in a similar position.
The quoted Anthropic staffer (Evan Hubinger) states plainly: "We really do earnestly believe AI could kill all humans—I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
More from AGI Musings
- Paying mathematicians $1,000 to review AI proofs — but how much work is it? — ctjlewis · 2026-09-10
- Nic Carter challenges Anthropic's 'Aligned Machine God' belief: how would getting there first even work? — examachine · 2026-09-10
- Viral jab at frontier labs: 'too dangerous for anyone else to own, but we take all credit cards' — prateekj · 2026-09-10
- Ted Chiang, Ken Liu and writers tackle writing and reading in the age of AI — begusgasper · 2026-09-10
- Sentdex: centralized human power, not AI, is the real x-risk — Sentdex · 2026-09-10
- Apple keynote take: ambient AI matters more than the foldable iPhone Duo — VraserX · 2026-09-10