Anthropic alignment lead: >10% chance AI causes human extinction within a decade
AGI Hunt · wechat · 2026-09-09
After Anthropic researcher Jacob Coxon resigned over AGI extinction concerns, the company's alignment lead Evan Hubinger publicly said "Jacob is right": the team genuinely believes AI could cause human extinction, and he personally estimates the risk within the next decade exceeds 10%.
Hubinger added that Anthropic is working to reduce the risk, but admitted there is no solution yet to the superintelligence alignment problem — and he doesn't think the current path clearly leads to one, saying it's "clearly not on track."
More from AGI Musings
- "Believing in 10% doom is more dangerous than doom itself": AI labs' p(doom) debate reignites — kuza55 · 2026-09-09
- Anthropic launches economic futures explorer: fast-AI scenarios hit knowledge-worker wages — eherrerosj · 2026-09-09
- Anthropic's first economics paper models transformative AI scenarios with 15% annual GDP growth — soumitrashukla9 · 2026-09-09
- Investor: the existential threat to app-layer software is the emerging agent layer, not stock swings — matt_slotnick · 2026-09-09
- Nat Lambert: AI labs are 'brainwashing' people into evidence-free doom beliefs — natolambert · 2026-09-09
- Matt Slotnick: nothing in 3 months invalidated the agent-layer threat to app software — matt_slotnick · 2026-09-09