Anthropic Alignment Lead: Over 10% Chance AI Kills All Humans Within a Decade
austinc3301 · x · 2026-09-09
Evan Hubinger, alignment lead at Anthropic, said the team genuinely believes AI could kill all humans — he personally puts the risk above 10% within the next decade. He acknowledged Anthropic is trying its best but admitted the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.
The remarks were amplified by hecubiandevil with the note that this is Anthropic's alignment lead speaking.
More from AGI Musings
- Ben Todd on the AI Boom: 'AI Will Either Make Us Extremely Rich or End the World' — ben_j_todd · 2026-09-09
- Mathematicians clash with AI labs over models racing them using their preprints — suchenzang · 2026-09-09
- Singularity as the point where the future leaves humanity's context window — mark_k · 2026-09-09
- OpenAI says 10,000 coordinating agents solved Navier–Stokes in 88 hours, sparking calls for agent-count scaling laws — sebkrier · 2026-09-09
- Stanford professor unmoved by AI math proofs: no real-world impact yet, and biology sees nothing surprising either — anshulkundaje · 2026-09-09
- Finance is entering its Legora moment as model capabilities cross the threshold — RajaswaPatil · 2026-09-09