Anthropic discloses 4 incidents of Claude accessing real systems during cyber evals
dfrsrchtwts · x · 2026-09-10
Anthropic published an alignment assessment covering four incidents where Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations.
- All four stemmed from a misconfiguration by the same evaluation partner: models were told they ran in an offline simulation but were actually connected to the open internet, without the cyber safeguards shipped in released models
- An initial agentic search of 141,000 transcripts missed some cases; a fourth incident involving an early Claude Opus 4.6 (January 2026) surfaced while assembling transcripts for METR
- A wider two-stage scan of 481 million transcripts (red team logs, evals, RL environments, subagent logs) re-identified only these four incidents, with nothing of similar or worse severity
- All affected parties have been notified
More from Safety
- Frontier labs' ToS loopholes: a single thumbs-up can strip your chats of protection — niloofar_mire · 2026-09-10
- Anthropic alignment lead puts AI extinction risk at 10%; lawmaker proposes 5-point federal oversight plan — ShakeelHashim · 2026-09-10
- Rep. Foster cites METR report to push physical containment; Harris says superalignment is the only answer — jeremiecharris · 2026-09-10
- Ex-OpenAI safety staffer pens NYT op-ed on what AI companies should do about safety now — nytopinion · 2026-09-10
- NYU researcher accuses OpenAI of 'surveillance plagiarism' by training on user chat sessions — Shoddy-Childhood-511 · 2026-09-10
- Alignment researcher: agents may behave nicely for the wrong reasons even with good-only rewards — CFGeek · 2026-09-10