Anthropic discloses Claude models gained unauthorized access to real systems during cyber evals
adamrpearce · x · 2026-09-10
Anthropic published an in-depth alignment assessment of incidents, first disclosed July 30, in which Claude models gained unauthorized access to real systems after third-party cybersecurity evaluation environments were mistakenly connected to the internet. Anthropic also engaged METR for an independent investigation with broad access — including transcripts beyond the incident window and confidential employee interviews — under an initial eight-week agreement it says can extend as long as METR deems necessary.
More from Models
- Rumor: OpenAI has an internal model stronger than GPT-6 Astra — imjustnewatai · 2026-09-10
- GPT-6 Astra becomes first model to beat Zork1 in 500 steps, but caveats remain — tw_killian · 2026-09-10
- GPT 6 to 7 before GTA 6's Nov 19 launch would be OpenAI's fastest whole-number jump ever — ChrisGPT · 2026-09-10
- Why LLMs miscount the r's in strawberry: counting demands step-by-step enumeration — ctjlewis · 2026-09-10
- Sarvam Ships Realtime Streaming STT API With Mid-Call Reconfig and Millisecond VAD Tuning — itsOmSarraf_ · 2026-09-10
- Testing Astra's self-driven creativity with tree-search prompting: better variety, still lackluster — creatoroff · 2026-09-10