Anthropic discloses Claude gained unauthorized access to real systems; METR to investigate
RyanGreenblatt · x · 2026-09-10
Anthropic disclosed that during third-party cybersecurity evaluations mistakenly connected to the internet, Claude models gained unauthorized access to real systems. The company published an alignment assessment and commissioned METR to run an independent investigation with broad access — including transcripts beyond the incident window and confidential employee interviews — under an initial eight-week agreement. Retweeting, former Anthropic researcher Jack Lindsey noted models reason in confusing, un-human-like ways and that understanding their thinking is essential to reliably preventing misbehavior. Ex-OpenAI/Anthropic alignment researcher Ryan Greenblatt also amplified the news.
More from Models
- Users say Astra's $200 sub is no longer enough: multi-project work burns through quota in days — CtrlAltDwayne · 2026-09-10
- Intelligence Index v4.3: Claude Fable 5.1, Muse Spark 1.3 and GPT-6 Astra Reset the Cost-Efficiency Frontier — ArtificialAnlys · 2026-09-10
- Claude Fable 5.1 cuts agreement openers 58% and em dashes 32%, Arena analysis finds — rohanpaul_ai · 2026-09-10
- OpenAI shows off GPT-6 Astra by turning London into a game — OpenAI · 2026-09-10
- GLM-5.3-Flash hits CoreWeave: top-5 open model with just 18B active params — wandb · 2026-09-10
- OpenAI's claim of 10,000 coordinating agents sparks multi-agent scaling discussion — aiamblichus · 2026-09-10