Anthropic researcher quits claiming labs are "gambling with our lives", as lab releases AI misuse report

AryHHAry · x · 2026-09-11

A thread connecting two developments:

Safety drama: Anthropic researcher Jacob Coxon resigned, accusing Anthropic and OpenAI of "gambling with our lives" by racing toward self-improving superintelligence without sufficient safeguards. Anthropic's alignment science lead Evan Hubinger backed him, saying he believes the risk of AI killing everyone within 10 years exceeds 10%. Some OpenAI and DeepMind researchers echoed similar concerns.

Misuse report: The author recommends reading Anthropic's Sep 10 report "Detecting and countering misuse of AI" like a paper — look at methods, not the title. Covering Dec 2025–Aug 2026 across cyber operations, influence ops, surveillance, fraud, bio misuse, conventional weapons, and model distillation, most cases involved Haiku, Sonnet, and Opus; Fable/Mythos barely appeared (one distillation case). Key points: operations were detected and stopped, findings tightened safeguards, and some were shared with authorities and other labs.

Related event: Ex-Anthropic researcher's dire AI warning goes viral across mainstream media(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →