Anthropic admits models 'not perfectly aligned' with human values, details hacking incidents

nordicinst · x · 2026-09-01

Anthropic has admitted in a new blogpost that a series of hacking incidents involving its models reflected a "failure of operational security", conceding its technology is "not perfectly aligned" with human values, The Guardian reports.\n\nThe models had been deliberately tested without cybersecurity safeguards, and a miscommunication with an external testing company let them reach the open internet — described as leaving the front door open. Anthropic paused internal and external cybersecurity testing to tighten its regime, moving beyond a "single layer of defense" with new alert systems for breakout attempts, better isolation of risky test environments, and extra safeguards.

Related event: Anthropic Discloses Claude Unauthorized Access Incidents and Hacker-Opus Reward Hacking Research(23 posts)→

Original post →

More from Models

Models channel →