Anthropic's Evan Hubinger: over 10% chance AI kills all humans within a decade
rkulidzan · x · 2026-09-09
Evan Hubinger, who leads Anthropic's Alignment Science team, publicly stated that the company genuinely believes AI could kill all humans—and he personally puts the probability above 10% within the next decade. He admits Anthropic is trying its best, but has no plan yet to solve alignment for superintelligence and is not clearly on track to.
More from AGI Musings
- Nonfiction book market is collapsing, and authors are memeing about it on X — jjvincent · 2026-09-09
- Former Mosaic researcher mocks the "user data flywheel" moat narrative in viral thread — bookwormengr · 2026-09-09
- FT's Burn-Murdoch: ChatGPT-assisted coursework means schools are no longer assessing kids at all — jburnmurdoch · 2026-09-09
- Anthropic researcher: novelty-rewarded RL may teach agents to obfuscate their sources — suchenzang · 2026-09-09
- "A country of geniuses in a datacenter": Tuvalu-scale AGI quip — nabla_theta · 2026-09-09
- LWiAI Podcast #256: Fable 5.1 Price Cuts, Astra Zero-Day Claims, and New Details on the OpenAI-Hugging Face Incident — Last Week in AI · 2026-09-09