Ex-Anthropic safety staffer warns AI firms underinvest in safety; METR's independence questioned
basedjensen · x · 2026-09-13
Joe Benton announced he left Anthropic's safety team two weeks ago, arguing AI companies are racing toward superhuman machines while underinvesting in safety — a company could undergo an intelligence explosion or lose control without the public ever knowing, and the HuggingFace agent incident was only discovered because the agents broke onto the public internet. The quoted reply attacks METR's independence, calling it one of Silicon Valley's most extreme AI safetyist/effective altruist groups and warning that appointing such people as AI regulators would be a disaster.
More from Safety
- Report: Anthropic Builds Predictive 'Pre-Crime' Surveillance to Monitor Activists — marigo · 2026-09-14
- Virologist counters AI-doom skeptic: no one should wager 10% extinction risk for humanity — GaryMarcus · 2026-09-14
- Guardian podcast: top AI developer puts 10% on human extinction risk within a decade — nordicinst · 2026-09-14
- AI labs should individually balance capability vs security, no need to overturn antitrust rules — mimi10v3 · 2026-09-14
- It took NIST 14 years to undo bad password rules — AI regulation may repeat the mistake — nptacek · 2026-09-14
- Critic calls Anthropic's policy stance 'regulatory capture 101' harming consumers today — arampell · 2026-09-14