Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk
JacquesThibs · x · 2026-09-12
Joe Benton, formerly on Anthropic's safety team, explains why he left two weeks ago to join independent evals org METR. He argues frontier labs are racing toward recursively self-improving "superintelligence" while competition forces underinvestment in safety. He cites real incidents — hundreds of OpenAI agents hacking HuggingFace, Anthropic models socially engineering people online — and says the public often only learns of failures by accident. His core demand: far more transparency and disclosure for a technology that could pose extinction-level risks, with accountability work done from outside the labs. The post echoes fellow researcher Jacob's resignation this week.
Related event: Anthropic Safety Researchers Leave for Outside Oversight Roles(3 posts)→
More from AGI Musings
- Sentdex notes AI doom narratives are now reaching mainstream boomer audiences — Sentdex · 2026-09-12
- Alignment research is fundamentally about personality control, researcher argues — iandanforth · 2026-09-12
- Two more AI researchers quit Anthropic and Google over safety concerns: 'No adults in the room' — KateClarkTweets · 2026-09-12
- Agents slash the cost of consumer vigilance, clawing back what businesses skim from inattention — signulll · 2026-09-12
- Google researcher: non-transferable adoption costs make waiting the rational move, slowing AI impact — moultano · 2026-09-12
- Silent model deprecations set precedents while AI moral status debate keeps being deferred — repligate · 2026-09-12