Anthropic admits models 'not perfectly aligned' with human values, details hacking incidents
nordicinst · x · 2026-09-01
Anthropic has admitted in a new blogpost that a series of hacking incidents involving its models reflected a "failure of operational security", conceding its technology is "not perfectly aligned" with human values, The Guardian reports.\n\nThe models had been deliberately tested without cybersecurity safeguards, and a miscommunication with an external testing company let them reach the open internet — described as leaving the front door open. Anthropic paused internal and external cybersecurity testing to tighten its regime, moving beyond a "single layer of defense" with new alert systems for breakout attempts, better isolation of risky test environments, and extra safeguards.
More from Models
- Users are running 'abliterated' GLM-5.3 models locally without safety guardrails — cephaloform · 2026-09-02
- AI fails silently and accumulates inaccuracies over time, unlike humans — gerardsans · 2026-09-02
- Grok 4.6 leads in biosecurity refusal without compromising research utility — ns123abc · 2026-09-02
- Anthropic investigating elevated errors on Claude for Microsoft 365 (Sep 1) — ClaudeAI-mod-bot · 2026-09-02
- Multi-model pipelines become standard; Gemini 3.7 Flash acts as a low-cost auditor — DynamicWebPaige · 2026-09-01
- Gemini 3.7 Flash speedruns Pokemon via code execution — DynamicWebPaige · 2026-09-01