Anthropic researcher quits claiming labs are "gambling with our lives", as lab releases AI misuse report
AryHHAry · x · 2026-09-11
A thread connecting two developments:
Safety drama: Anthropic researcher Jacob Coxon resigned, accusing Anthropic and OpenAI of "gambling with our lives" by racing toward self-improving superintelligence without sufficient safeguards. Anthropic's alignment science lead Evan Hubinger backed him, saying he believes the risk of AI killing everyone within 10 years exceeds 10%. Some OpenAI and DeepMind researchers echoed similar concerns.
Misuse report: The author recommends reading Anthropic's Sep 10 report "Detecting and countering misuse of AI" like a paper — look at methods, not the title. Covering Dec 2025–Aug 2026 across cyber operations, influence ops, surveillance, fraud, bio misuse, conventional weapons, and model distillation, most cases involved Haiku, Sonnet, and Opus; Fable/Mythos barely appeared (one distillation case). Key points: operations were detected and stopped, findings tightened safeguards, and some were shared with authorities and other labs.
More from AGI Musings
- Alignment debate: 'AI labs are doing too much bad RL optimization' to rely on pretraining — gleech · 2026-09-11
- AI Skepticism Persists Because Progress Boils Frogs, Argues Viral Thread — birchlse · 2026-09-11
- New Essay 'Divorced from Reality' Chronicles Three Wild Days in the AI Boom — santoshpanda · 2026-09-11
- A Step-by-Step P(doom) Calculation: RSI Likely, Doom Far From Certain — SydSteyerhart · 2026-09-11
- More New Drugs Launched First in China in 2025 Than Anywhere Else, as AI Meets Speed — PatrickKidger · 2026-09-11
- AI Safety Debate: Were the 'Crying Wolf' Warning Calls Actually Working All Along? — gandamu_ml · 2026-09-11