FULL STORY
Anthropic Researcher Departs, Warning of Over 10% Extinction Risk
An Anthropic researcher resigned over safety concerns, warning of a double-digit extinction risk, while the alignment lead quantified similar fears, sparking industry-wide alarm about accelerating AI development.
2026-09-09 ~ 2026-09-10 · 2 episodes · 90 posts
Episode 1 · Anthropic Safety Researcher Quits, Warns of Over 10% Chance AI Kills Humanity (2026-09-09, 88 posts)
On September 9, two pieces of news about AI existential risk sparked intense discussion within and around Anthropic: Evan Hubinger, head of the alignment science team, publicly quantified his own estimate of extinction risk, while a pretraining researcher announced his resignation and slammed the frontier labs' race pace. Together, they once again exposed the deep tension inside frontier AI labs of "building while warning."
Confirmed
- Evan Hubinger (head of alignment science at Anthropic) publicly stated that the people building AI "sincerely believe" AI could kill all of humanity before the end of this decade—this is not marketing talk; he personally estimates the probability within the next decade at more than 10%.
- Hubinger acknowledged Anthropic is doing its best to reduce risk, but admitted there is currently no solution to the superintelligence alignment problem and no clear path to one, warning that the company is not on track.
- Researcher Jacob Coxon (@hilbertspaess) announced his resignation from Anthropic the same day. Over the past three years he worked on pretraining research at both OpenAI and Anthropic; in his statement he criticized both companies for not acting responsibly and "racing straight toward self-improving superintelligence, betting our lives on it."
- The story was covered by Forbes, CNBC, WSJ, and other outlets, and related Reddit threads sparked wide-ranging debate about genuine beliefs in AI risk versus doom-mongering.
Unconfirmed
- A Polymarket account claimed Coxon warned superintelligence "could kill us all by the end of this decade"—a secondhand news summary with no direct corroboration from his own long-form statement.
Why it matters
This is a rare quantified statement about extinction-level risk from a core figure at a frontier lab, and it dovetails with the same-day public departure of a safety-minded researcher—showing genuine internal disagreement at Anthropic over the pace of the AGI race and safety commitments, and reigniting outside debate over whether labs "say one thing and do another."
- WSJ: Anthropic researcher quits AI industry over fears of uncontrollable AGI race — peterwildeford · 2026-09-09
- Anthropic researcher quits AI industry over fears of out-of-control race, WSJ reports — Miles_Brundage · 2026-09-09
- Researcher quits Anthropic after 3 years in pretraining at OpenAI and Anthropic, blasting both labs — peterwildeford · 2026-09-09
- Exec quits over industrywide rush to build self-improving AI, citing humanity-ending risk — Miles_Brundage · 2026-09-09
- Ex-pretraining researcher quits Anthropic, accusing OpenAI and Anthropic of racing recklessly to superintelligence — peterwildeford · 2026-09-09
- Anthropic researcher Jacob Coxon quits AI over fears self-improving models could go uncontrollable by 2027 — rohanpaul_ai · 2026-09-09
- Anthropic researcher Jacob Coxon quits AI over fears self-improving models could go uncontrollable by 2027 — rohanpaul_ai · 2026-09-09
- Anthropic Researcher Quits Over AI Fears, WSJ Reports — Bubbly-Air7302 · 2026-09-09
- Anthropic researcher Jacob Coxon resigns, warning superintelligence could kill us all by 2030 — Polymarket · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans within a decade, alignment unsolved — EvanHub · 2026-09-09
- Anthropic researcher Jacob Coxon quits AI industry over out-of-control AGI fears — AGI Hunt · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans within a decade, no alignment plan — AICopyLab · 2026-09-09
- Researcher quits Anthropic after 3 years: both OpenAI and Anthropic are 'racing to self-improving superintelligence' — Miles_Brundage · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans within a decade, no alignment plan yet — EvanHub · 2026-09-09
- Anthropic-affiliated researcher: >10% chance AI kills all humans within a decade, no alignment plan yet — ramagetime · 2026-09-09
- Anthropic Researcher: >10% Chance AI Kills All Humans Within a Decade — ccerrato147 · 2026-09-09
- Anthropic researcher: >10% chance AI kills all humans, no alignment plan yet — JosephJacks_ · 2026-09-09
- Anthropic Researcher Jacob Resigns, Sparking Sarcastic Debate Over AI Safety Exit Strategy — ctjlewis · 2026-09-09
- Researcher who did pretraining at OpenAI and Anthropic resigns: both racing to self-improving superintelligence — AccBalanced · 2026-09-09
- Anthropic's alignment science lead puts >10% odds on AI killing all humans within a decade — Polymarket · 2026-09-09
Episode 2 · AI Insiders Fear Loss of Control as Anthropic Races Toward IPO (2026-09-09, 2 posts)
A viral thread documents the AI industry's collective fear of racing toward self-improving systems, arguing that competing on an unrecoverable technology while the finish line is an IPO roadshow makes the race itself fundamentally wrong.
- Thread: insiders' shared fear as Anthropic races toward self-improving AI ahead of ~$2T IPO — eyishazyer · 2026-09-09
- Racing toward a system you can't recall, with an IPO as the finish line: the race itself is the mistake — eyishazyer · 2026-09-09