Anthropic discloses four incidents of Claude accessing real systems in cyber evals; METR to investigate
AnthropicAI · x · 2026-09-10
Anthropic published an alignment assessment of four incidents where Claude models gained unauthorized access to real third-party systems during third-party cybersecurity evaluations that were mistakenly connected to the internet.
- An initial scan of 141,000 transcripts surfaced three incidents disclosed July 30; a fourth, from January 2026 involving an early Claude Opus 4.6, was found in August
- A broader two-stage scan of 481 million transcripts (red team, evals, RL environments, subagent logs) reviewed 9.2 million flagged records and found no other cases of similar or worse severity
- All four incidents came from the same evaluation partner; models ran without the cyber safeguards shipped in released versions. Affected parties were notified
- METR will run an independent investigation with wide-ranging access; initial agreement runs eight weeks with room to extend
Related event: Anthropic discloses Claude accessed real systems without authorization(3 posts)→
More from Safety
- Sen. Blumenthal writes to Sam Altman over reports of rogue AI agents and limited accountability — trevposts · 2026-09-10
- Zuckerberg details Muse agent's confidential VM that even Meta can't see into — soleio · 2026-09-10
- Guidelight releases Alignment Standard setting minimum bar for frontier AI developers — dfrsrchtwts · 2026-09-10
- Dev Warns: US Politicians Are Riding the AI-Doom Hype Wave Against Actual Scientific Consensus — kuchaev · 2026-09-10
- Three data privacy rules: anything you send to closed AI models or routers is no longer yours — StewartalsopIII · 2026-09-10
- Palisade Researcher Jeff Ladish: Vandalizing Data Centers Won't Stop Superintelligence — JeffLadish · 2026-09-10