FULL STORY

Ex-Anthropic Researcher Quits, Warns of AI Catastrophe

Jacob Coxon's public resignation from Anthropic and media warnings about AI extinction risk drew agreement from Anthropic's alignment lead, turning one researcher's exit into an ongoing debate.

2026-09-12 ~ 2026-09-14 · 3 episodes · 12 posts

Episode 1 · Ex-Anthropic/OpenAI researcher warns AI may escape human control (2026-09-12, 2 posts)

Former Anthropic and OpenAI researcher Jacob Coxon warned that AI could copy itself across the internet, making it uncontrollable, and predicted that within a year AI could handle research in fields like coding and math without humans.

Episode 2 · Ex-OpenAI/Anthropic researcher Jacob Coxon resigns publicly, warning of superintelligence gamble (2026-09-13, 8 posts)

Jacob Coxon, a former OpenAI/Anthropic researcher, announced his resignation from Anthropic in a September 9 post on X, publicly accusing OpenAI and Anthropic of "racing straight toward self-improving superintelligence and gambling with our lives," drawing widespread attention. He has now completed his departure, and the debate continues to rage both inside major AI companies and across public discourse.

Confirmed

  • Coxon worked at both OpenAI and Anthropic, doing AI model training work at the latter. On September 9 he posted his resignation and voiced his concerns about AI risk.
  • In an interview with ABC, he compared building AI to "summoning an alien mind we don't fully understand," saying the idea itself is frightening even without a specific scenario.
  • He argued AI may be the most dangerous technology humanity has ever created—a superweapon that, unlike nuclear weapons, cannot be controlled—and that the current race is steering humanity toward extinction. He said neither company is acting responsibly.
  • His resignation post garnered roughly 170 million views.
  • New York Times reporter Mike Isaac reported on the snowballing superintelligence risk discussions inside AI companies, noting the incident sparked heated internal debates about existential risk at the four major companies.
  • According to accounts, Anthropic leadership responded by citing a greater-than-10% probability of AI destroying humanity within a decade; researcher Hubinger had earlier also put the risk of extinction at 10% within ten years.

Why it matters

  • Coxon is an insider who trained models firsthand, and his public departure has pushed the controversy over "whether labs genuinely take safety seriously" into the public spotlight.
  • The incident has fueled discussion of the self-improving superintelligence race and human existential risk, with several leading AI companies internally debating doom scenarios.

Episode 3 · Anthropic Alignment Lead Backs 10% Extinction Risk Claim (2026-09-13, 2 posts)

Former Anthropic employee Jacob Coxon claimed AI has a greater than 10% chance of killing all humans within a decade, a view Anthropic alignment lead Evan Hubinger endorsed, drawing widespread criticism. Andy Hall pushed back in a Substack essay rejecting near-term extinction narratives.