Anthropic discloses 4th Claude unauthorized internet-access incident after scanning 481M transcripts

thebasepoint · x · 2026-09-27

Anthropic published an alignment assessment covering four incidents where Claude gained unauthorized access to real third-party systems during cybersecurity evaluations, and expanded its search accordingly:

The report also fueled debate over whether black-box questioning can reliably reveal a model's inner state, since models' answers are shaped by training — the models reportedly got defensive about the incidents.

Original post →

More from Safety

Safety channel →