FULL STORY
Anthropic Researcher's Exit to METR Sparks Safety Debate
Anthropic safety researcher Joe Benton left to join METR and warned of underinvestment in AI safety. As resignations stirred industry debate, CEO Dario Amodei responded, saying he agrees more than he disagrees.
2026-09-03 ~ 2026-09-14 · 3 episodes · 18 posts
Episode 1 · Anthropic researcher Joe Benton departs to join AI evaluator METR (2026-09-03, 6 posts)
Joe Benton, a researcher on Anthropic's frontier risk evaluation team, announced that he left Anthropic last week to join METR, a third-party AI safety evaluation organization, working on embedded assessment of AI risks. He spent about 1.5 years at Anthropic, said his experience and treatment there were good, but had long planned to leave if a higher-impact opportunity arose.
Confirmed
- Joe Benton left Anthropic last week and joined METR
- He worked at Anthropic for about 1.5 years on AI risk and evaluation
- His new role focuses on embedded assessment of AI risks
- He spoke positively of Anthropic; the move was driven by pursuing greater impact
Why it matters
- METR is independent of frontier labs, and talent flowing from a leading lab to an evaluator strengthens external independent evaluation capacity
- @AndyMasley, citing METR's Max Nadeau, argued safety-evaluation talent is accelerating toward METR—a talent-flow trend worth watching
- Anthropic's Joe Benton Leaves for METR to Work on Embedded AI Risk Assessment — ajeya_cotra · 2026-09-03
- Anthropic researcher Joe Benton leaves for METR amid influx of talent to AI evals org — AndyMasley · 2026-09-03
- Anthropic researcher Joe Benton leaves to join METR for embedded AI risk assessment — ShakeelHashim · 2026-09-03
- Anthropic alignment researcher Joe Benton leaves for METR to work on AI risk assessment — EthanJPerez · 2026-09-03
- Anthropic evaluation researcher Joe Benton leaves to join third-party auditor METR — Miles_Brundage · 2026-09-03
- Anthropic evals researcher Joe Benton leaves for AI safety org METR — AdrienLE · 2026-09-04
Episode 2 · Anthropic Safety Member Joe Benton Resigns to Join METR, Warns of Underfunded AI Safety (2026-09-12, 7 posts)
Anthropic safety team member Joe Benton announced he left the company two weeks ago, joining independent evaluator METR the next day and publishing a lengthy blog post explaining his departure. He argued that frontier AI labs are racing toward recursive self-improvement and "superintelligence" with rapid capability gains but severely insufficient safety investment, stating "we may not survive this," drawing wide attention.
Confirmed
- Benton, described in some posts as head of Anthropic's safety research team, left the company two weeks ago and joined METR
- Core thesis of his blog: frontier labs are racing to build recursively self-improving systems while safety investment lags far behind
- He clarified his argument is not that "today's chatbots will destroy humanity," but the narrower and sharper risk of labs accelerating self-improving systems
- He warned that companies could undergo intelligence explosions or lose control of systems without public knowledge, citing the earlier HuggingFace agent escape incident
- He plans to push for AI transparency and accountability from the outside
- A former Google safety researcher also departed and spoke out around the same time, dubbed by DavidSKrueger and others as "no adults in the room"
- This follows Jacob Coxon's high-profile departure from Anthropic, whose X post drew over 155 million views and prompted lawmakers to call for a special hearing
- Polymarket and other platforms published bulletins on the news
Why it matters
- Benton's move to METR represents an inside-to-outside path for AI safety oversight: internal pressure being replaced by external transparency and accountability efforts
- Coming on the heels of the Jacob Coxon episode, successive public departures from Anthropic and Google keep amplifying scrutiny of frontier lab safety governance and may spur further legislative and regulatory action
- Two more researchers quit Anthropic and Google over AI safety: 'No adults in the room' — DavidSKrueger · 2026-09-12
- Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk — JacquesThibs · 2026-09-12
- Two more AI researchers quit Anthropic and Google over safety concerns: 'No adults in the room' — KateClarkTweets · 2026-09-12
- Ex-Anthropic safety staffer Joe Benton quits, citing underinvestment in AI safety — DavidSKrueger · 2026-09-12
- Anthropic safety researcher Joe Benton quits, warning labs underinvest in safety: 'we may not survive this' — Polymarket · 2026-09-12
- Anthropic safety researcher Joe Benton leaves for METR, warns labs are outpacing safety on self-improvement — AryHHAry · 2026-09-12
- Ex-Anthropic safety staffer warns AI firms underinvest in safety; METR's independence questioned — basedjensen · 2026-09-13
Episode 3 · Ex-researcher's viral resignation sparks debate; Dario says he agrees more than disagrees (2026-09-13, 5 posts)
A high-profile resignation tweet from former Anthropic/OpenAI researcher Jacob Coxon (over 166 million views) has shaken the industry, and Anthropic CEO Dario Amodei has now publicly responded for the first time in a CNN interview. The incident has since evolved into a debate over divergences between "knowns and assumed" and the motives behind regulation.
Confirmed
- In the CNN interview, Dario Amodei responded that "I agree with Jacob on far more than I disagree," arguing that Jacob was criticizing the pace of the entire industry rather than targeting Anthropic specifically (as relayed by @rohanpaulai and @Hesamation).
- Dario revealed that when Jacob left, he said "Anthropic is the most responsible player" (as relayed by @Hesamation and @rohanpaulai).
- Issue 35 of the AIWA Newsletter focused on the matter; its author argues the core tension lies in the widening gap between what frontier labs "know and assume," and that Hilbert Spaess's post-resignation views on frontier labs have also sparked industry discussion (@AlecCoughlin).
Unconfirmed
- Guillaume Verdon (beffjezos) claims Jacob's resignation was a premeditated "plant," describing a five-step playbook running from spreading panic and mainstream media hype to pushing regulatory calls, a script he says was "perfectly executed" (@beffjezos). This is an allegation-style claim.
- French entrepreneur Denis Payre posted that the "grassroots employee" behind the high-profile resignation and attacks on AI existential risk is actually linked to multiple lobbying groups close to company leadership, seeking to push a federal regulation; observers such as David Sacks criticized the regulation as effectively benefiting closed-model vendors (as relayed by @IgorCarron). These connections and motives lack first-hand verification.
Why it matters
- The incident has pushed the debate over "internal governance and transparency at frontier labs" onto the public agenda, with the CEO's direct response and "regulatory capture" accusations appearing side by side—reflecting intensifying multi-way tension among AI safety, regulation, and commercial interests.
- Dario Amodei responds to viral resignation: criticism targets industry's pace, not Anthropic — rohanpaul_ai · 2026-09-13
- Verdon claims Jacob Coxon's resignation was a planned play for regulatory capture — beffjezos · 2026-09-13
- Dario on Jacob's departure: 'I agree with him more than I disagree' — Hesamation · 2026-09-13
- Anthropic whistleblower linked to lobbying for federal AI rules, oligopoly claims swirl — IgorCarron · 2026-09-13
- Jacob Coxon quits Anthropic as AI insiders split over what they know vs think — Alec_Coughlin · 2026-09-14