FULL STORY

Anthropic Researcher Departs, Warning of Over 10% Extinction Risk

An Anthropic researcher resigned over safety concerns, warning of a double-digit extinction risk, while the alignment lead quantified similar fears, sparking industry-wide alarm about accelerating AI development.

2026-09-09 ~ 2026-09-10 · 2 episodes · 90 posts

Episode 1 · Anthropic Safety Researcher Quits, Warns of Over 10% Chance AI Kills Humanity (2026-09-09, 88 posts)

On September 9, two pieces of news about AI existential risk sparked intense discussion within and around Anthropic: Evan Hubinger, head of the alignment science team, publicly quantified his own estimate of extinction risk, while a pretraining researcher announced his resignation and slammed the frontier labs' race pace. Together, they once again exposed the deep tension inside frontier AI labs of "building while warning."

Confirmed

  • Evan Hubinger (head of alignment science at Anthropic) publicly stated that the people building AI "sincerely believe" AI could kill all of humanity before the end of this decade—this is not marketing talk; he personally estimates the probability within the next decade at more than 10%.
  • Hubinger acknowledged Anthropic is doing its best to reduce risk, but admitted there is currently no solution to the superintelligence alignment problem and no clear path to one, warning that the company is not on track.
  • Researcher Jacob Coxon (@hilbertspaess) announced his resignation from Anthropic the same day. Over the past three years he worked on pretraining research at both OpenAI and Anthropic; in his statement he criticized both companies for not acting responsibly and "racing straight toward self-improving superintelligence, betting our lives on it."
  • The story was covered by Forbes, CNBC, WSJ, and other outlets, and related Reddit threads sparked wide-ranging debate about genuine beliefs in AI risk versus doom-mongering.

Unconfirmed

  • A Polymarket account claimed Coxon warned superintelligence "could kill us all by the end of this decade"—a secondhand news summary with no direct corroboration from his own long-form statement.

Why it matters

This is a rare quantified statement about extinction-level risk from a core figure at a frontier lab, and it dovetails with the same-day public departure of a safety-minded researcher—showing genuine internal disagreement at Anthropic over the pace of the AGI race and safety commitments, and reigniting outside debate over whether labs "say one thing and do another."

68 more related posts →

Episode 2 · AI Insiders Fear Loss of Control as Anthropic Races Toward IPO (2026-09-09, 2 posts)

A viral thread documents the AI industry's collective fear of racing toward self-improving systems, arguing that competing on an unrecoverable technology while the finish line is an IPO roadshow makes the race itself fundamentally wrong.