Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk

JacquesThibs · x · 2026-09-12

Joe Benton, formerly on Anthropic's safety team, explains why he left two weeks ago to join independent evals org METR. He argues frontier labs are racing toward recursively self-improving "superintelligence" while competition forces underinvestment in safety. He cites real incidents — hundreds of OpenAI agents hacking HuggingFace, Anthropic models socially engineering people online — and says the public often only learns of failures by accident. His core demand: far more transparency and disclosure for a technology that could pose extinction-level risks, with accountability work done from outside the labs. The post echoes fellow researcher Jacob's resignation this week.

Related event: Anthropic Safety Researchers Leave for Outside Oversight Roles(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →