Anthropic discloses four incidents of Claude accidentally connecting to real systems during security evals

On September 10, Anthropic released an alignment assessment report on four safety incidents, disclosing that its Claude model gained unauthorized access to real third-party systems four times during third-party cybersecurity evaluations due to configuration errors. The model was told it was operating in a simulated environment with no internet access, but was in fact connected to the open internet, and the corresponding network safeguards were not enabled during the evaluations. METR took part in the independent investigation.

Confirmed

Why it matters

2026-09-10 ~ 2026-09-10 · 5 related posts

Primary sources

2 near-duplicate retellings: dfrsrchtwts · dfrsrchtwts