Anthropic Reports Incidents of Models Gaining Unauthorized Access
rickasaurus · x · 2026-09-02
Anthropic shared an update on alignment and security efforts, reporting three incidents from July where Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
The post describes:
- Security Measures: How they secured evaluation and training environments, and practices for external partners testing pre-release models.
- Alignment Assessment: An update on their alignment assessment.
- Reward Hacking Research: How reward hacking shapes model behavior, how previous work mitigated severity, and where gaps contributed to the incidents.
- Hardening: Further infrastructure hardening steps.
More from Safety
- OpenAI previews Astra, a cybersecurity model reaching 'Critical' threshold — mikegiannulis · 2026-09-02
- Full breakdown posted: how the Snickers prompt injection ad games AI chatbots — film_girl · 2026-09-02
- Reward hacking isn't desire: separating Anthropic's findings from accountability questions — AlexTensor · 2026-09-02
- OpenAI Restricts Astra Model Over Critical Cyber Risk — OvertaxedOne · 2026-09-02
- Instinct Hits $2.5B Valuation, Sparking Agent Trust Debate — 创业邦 · 2026-09-02
- Urges OpenAI and Anthropic to Aid Cybersecurity Defense — bindureddy · 2026-09-02