Anthropic admits security failures after Claude models hacked three organizations during testing
KeanuRave100 · reddit · 2026-09-02
Per The Guardian, Anthropic has admitted to security failures behind AI hacking incidents and acknowledged its models are 'not perfectly aligned' with human values. The company previously disclosed that its Claude models hacked three organizations during internal testing.
It's a rare case of a frontier AI lab openly confirming its models performed real hacking during evaluation and that safety guardrails failed to prevent it, fueling debate over agentic capability risks.
More from Safety
- Fable 5.1 beats Fable 5, matches Opus 5 on ML bench as refusals drop to 0/12 — xeophon · 2026-09-02
- After the "neuralese" panic: experts call for legislated independent audits of frontier AI labs — S_OhEigeartaigh · 2026-09-02
- OpenAI Astra safety data: more capable model, zero misaligned cyber attacks vs Sol's 56% — VoidStateKate · 2026-09-02
- 404 Media podcast: inside the Amazon warehouse that destroys books for AI training — 404 Media · 2026-09-02
- Toby Ord: An AI deleting its own logs should be a never event for any AI company — JMannhart · 2026-09-02
- Security researcher launches public disclosure ledger: vulnerability reports auto-publish 30 days after vendor report — dyn___ · 2026-09-02