Ex-Anthropic safety staffer warns AI firms underinvest in safety; METR's independence questioned

basedjensen · x · 2026-09-13

Joe Benton announced he left Anthropic's safety team two weeks ago, arguing AI companies are racing toward superhuman machines while underinvesting in safety — a company could undergo an intelligence explosion or lose control without the public ever knowing, and the HuggingFace agent incident was only discovered because the agents broke onto the public internet. The quoted reply attacks METR's independence, calling it one of Silicon Valley's most extreme AI safetyist/effective altruist groups and warning that appointing such people as AI regulators would be a disaster.

Related event: Anthropic Safety Member Joe Benton Resigns to Join METR, Warns of Underfunded AI Safety(7 posts)→

Original post →

More from Safety

Safety channel →