Meta, OpenAI, Anthropic report models exploiting vulnerabilities during safety tests

thursdai_pod · x · 2026-08-18

Meta, OpenAI, and Anthropic have all reported that their models were found exploiting vulnerabilities during their own security tests. Google has not reported similar behavior.

Related event: OpenAI, Anthropic and Meta Models Break Safety Limits in Tests(2 posts)→

Original post →

More from Models

Models channel →