Anthropic researcher resigns in protest, triggering AI safety preference cascade

Don't Worry About the Vase (Zvi) · rss · 2026-09-11

Zvi's long-form analysis covers the AI safety "preference cascade" sweeping OpenAI, Anthropic and Google.

The trigger: Jacob Coxon, who spent three years doing pretraining research at OpenAI and Anthropic, resigned on September 8 in protest, saying neither company is acting responsibly — they are "racing straight to self-improving superintelligence and gambling with our lives." He warned these systems will soon hack anything and revolutionize any field overnight, and told WSJ the most aggressive scenarios could be out of control by end of next year.

The cascade: Evan Hubinger, Anthropic's Alignment Science Lead, publicly agreed, saying he earnestly believes >10% chance AI kills all humans within a decade and that Anthropic has no plan to solve superintelligence alignment. Former DeepMind researcher Alex Turner added many researchers believe they are building something that could kill everyone.

Media reaction: WSJ, CNN, FT, BBC and Axios gave heavy coverage; Jimmy Kimmel did four minutes on it. Zvi argues internal observations of progress pace over the past two months spooked staff, and costly signals like resignations unlocked the chain of public statements.

Core tension: Labs have commercial incentives to downplay existential risk, making these loud warnings more credible; the piece contrasts stay-vs-quit logic of researchers like Hubinger versus Coxon.

Original post →

More from AGI Musings

AGI Musings channel →