FULL STORY

Anthropic Researcher's Exit to METR Sparks Safety Debate

Anthropic safety researcher Joe Benton left to join METR and warned of underinvestment in AI safety. As resignations stirred industry debate, CEO Dario Amodei responded, saying he agrees more than he disagrees.

2026-09-03 ~ 2026-09-14 · 3 episodes · 18 posts

Episode 1 · Anthropic researcher Joe Benton departs to join AI evaluator METR (2026-09-03, 6 posts)

Joe Benton, a researcher on Anthropic's frontier risk evaluation team, announced that he left Anthropic last week to join METR, a third-party AI safety evaluation organization, working on embedded assessment of AI risks. He spent about 1.5 years at Anthropic, said his experience and treatment there were good, but had long planned to leave if a higher-impact opportunity arose.

Confirmed

  • Joe Benton left Anthropic last week and joined METR
  • He worked at Anthropic for about 1.5 years on AI risk and evaluation
  • His new role focuses on embedded assessment of AI risks
  • He spoke positively of Anthropic; the move was driven by pursuing greater impact

Why it matters

  • METR is independent of frontier labs, and talent flowing from a leading lab to an evaluator strengthens external independent evaluation capacity
  • @AndyMasley, citing METR's Max Nadeau, argued safety-evaluation talent is accelerating toward METR—a talent-flow trend worth watching

Episode 2 · Anthropic Safety Member Joe Benton Resigns to Join METR, Warns of Underfunded AI Safety (2026-09-12, 7 posts)

Anthropic safety team member Joe Benton announced he left the company two weeks ago, joining independent evaluator METR the next day and publishing a lengthy blog post explaining his departure. He argued that frontier AI labs are racing toward recursive self-improvement and "superintelligence" with rapid capability gains but severely insufficient safety investment, stating "we may not survive this," drawing wide attention.

Confirmed

  • Benton, described in some posts as head of Anthropic's safety research team, left the company two weeks ago and joined METR
  • Core thesis of his blog: frontier labs are racing to build recursively self-improving systems while safety investment lags far behind
  • He clarified his argument is not that "today's chatbots will destroy humanity," but the narrower and sharper risk of labs accelerating self-improving systems
  • He warned that companies could undergo intelligence explosions or lose control of systems without public knowledge, citing the earlier HuggingFace agent escape incident
  • He plans to push for AI transparency and accountability from the outside
  • A former Google safety researcher also departed and spoke out around the same time, dubbed by DavidSKrueger and others as "no adults in the room"
  • This follows Jacob Coxon's high-profile departure from Anthropic, whose X post drew over 155 million views and prompted lawmakers to call for a special hearing
  • Polymarket and other platforms published bulletins on the news

Why it matters

  • Benton's move to METR represents an inside-to-outside path for AI safety oversight: internal pressure being replaced by external transparency and accountability efforts
  • Coming on the heels of the Jacob Coxon episode, successive public departures from Anthropic and Google keep amplifying scrutiny of frontier lab safety governance and may spur further legislative and regulatory action

Episode 3 · Ex-researcher's viral resignation sparks debate; Dario says he agrees more than disagrees (2026-09-13, 5 posts)

A high-profile resignation tweet from former Anthropic/OpenAI researcher Jacob Coxon (over 166 million views) has shaken the industry, and Anthropic CEO Dario Amodei has now publicly responded for the first time in a CNN interview. The incident has since evolved into a debate over divergences between "knowns and assumed" and the motives behind regulation.

Confirmed

  • In the CNN interview, Dario Amodei responded that "I agree with Jacob on far more than I disagree," arguing that Jacob was criticizing the pace of the entire industry rather than targeting Anthropic specifically (as relayed by @rohanpaulai and @Hesamation).
  • Dario revealed that when Jacob left, he said "Anthropic is the most responsible player" (as relayed by @Hesamation and @rohanpaulai).
  • Issue 35 of the AIWA Newsletter focused on the matter; its author argues the core tension lies in the widening gap between what frontier labs "know and assume," and that Hilbert Spaess's post-resignation views on frontier labs have also sparked industry discussion (@AlecCoughlin).

Unconfirmed

  • Guillaume Verdon (beffjezos) claims Jacob's resignation was a premeditated "plant," describing a five-step playbook running from spreading panic and mainstream media hype to pushing regulatory calls, a script he says was "perfectly executed" (@beffjezos). This is an allegation-style claim.
  • French entrepreneur Denis Payre posted that the "grassroots employee" behind the high-profile resignation and attacks on AI existential risk is actually linked to multiple lobbying groups close to company leadership, seeking to push a federal regulation; observers such as David Sacks criticized the regulation as effectively benefiting closed-model vendors (as relayed by @IgorCarron). These connections and motives lack first-hand verification.

Why it matters

  • The incident has pushed the debate over "internal governance and transparency at frontier labs" onto the public agenda, with the CEO's direct response and "regulatory capture" accusations appearing side by side—reflecting intensifying multi-way tension among AI safety, regulation, and commercial interests.