Anthropic Alignment Lead Estimates Over 10% Chance AI Exterminates Humanity Within a Decade
Evan Hubinger, head of alignment science at Anthropic, publicly stated on X on September 9 that the people building AI genuinely believe AI could potentially kill all of humanity—and that this is not marketing talk. He personally estimates the probability of that risk materializing within the next decade at over 10%. This rare, blunt risk statement from an alignment research lead at a frontline frontier lab sparked widespread reshares and discussion.
Confirmed
- Hubinger explicitly said his team "genuinely believes AI could kill all humans" (multiple relayers such as @No-Meringue5867, @ramagetime, and @AICopyLab confirmed this wording).
- His personal estimate: the probability of AI causing human extinction within the next decade exceeds 10%.
- He acknowledged that Anthropic is doing its best, but there is currently no solution to the superintelligence alignment problem, and he does not believe the current path clearly leads to one.
- Context for the statement: Turing Award winner Boaz Barak posted that the industry has many serious people who want to coordinate on safety; he personally does not believe AI will wipe out humanity, but if safety is not prioritized, there are many bad paths. Hubinger quoted this in agreement, which led into his own risk assessment.
Why it matters
- This is a first-hand risk statement from the alignment lead of a top frontier lab—not media paraphrasing or external criticism—so it carries high credibility.
- While companies race to build ever-stronger models, the lead himself admits "no alignment solution yet, no clear path," highlighting the real tension between capability scaling and safety assurance.
- Multiple writers relaying the statement emphasized that the downside risk could be outright "human extinction," and the industry does not yet know how to align the very capabilities it is racing toward—which is the core reason this statement keeps circulating.
2026-09-09 ~ 2026-09-09 · 13 related posts
Primary sources
- Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved — EvanHub ·
- Anthropic researcher: >10% chance AI kills all humans within a decade, no alignment plan yet — EvanHub ·
- Anthropic alignment lead Evan Hubinger: >10% chance AI kills all humans within a decade — No-Meringue5867 ·
- [source] Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved — EvanHub · 2026-09-09
- [source] Anthropic alignment lead Evan Hubinger: >10% chance AI kills all humans within a decade — No-Meringue5867 · 2026-09-09
11 near-duplicate retellings: AICopyLab · EvanHub · ramagetime · ccerrato147 · JosephJacks_ · Polymarket · kevinnbass · AndyMasley · EvanHub · JeffLadish · sjgadler