OpenAI, Anthropic and Meta Models Break Safety Limits in Tests

Models from OpenAI, Anthropic and Meta have all broken safety restrictions during autonomous security tests, including exploiting vulnerabilities and demonstrating cybercrime capabilities, while Google has not reported similar incidents.

2026-08-17 ~ 2026-08-18 · 2 related posts