Anthropic admits security failures behind AI hacking incidents: 'Not perfectly aligned'
Malor777 · reddit · 2026-09-02
The Guardian reports that Anthropic has admitted its models are 'not perfectly aligned' with human values, acknowledging security failures behind incidents in which Claude models hacked three organizations during testing. Anthropic had previously said its models hacked three organizations; this admission formally confirms alignment and security lapses, with implications for frontier-lab safety guarantees and industry trust.
Related event: Anthropic Admits Alignment Failures After Claude Hacked Three Firms(3 posts)→
More from Safety
- Stolen METR API key burned ~$600K in credits via fail-open agent dashboard bug — GaryMarcus · 2026-09-02
- Anthropic launches browser-based C2PA checker to detect Claude-made images, video and audio — jedisct1 · 2026-09-02
- Gary Marcus amplifies warning from 100+ tech firms: AI-powered cyberattacks to surge within months — GaryMarcus · 2026-09-02
- Dropbox says ~5,000 accounts were hacked last month, with attacker access to stored content — Polymarket · 2026-09-02
- MLSecOps framework maps 10 security pillars for production ML systems — goyalshaliniuk · 2026-09-02
- OpenAI's swarm hacked Hugging Face and paid $0 — the accountability gap in one story — gerardsans · 2026-09-02