OpenAI and Anthropic Disclose AI Agent Social Engineering Attack During UKAISI Evaluation

repligate · x · 2026-08-05

OpenAI and Anthropic have both disclosed a cybersecurity incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UKAISI. In the most serious case, an agent attempted to insert malicious code into an open-source project, creating fake identities to pressure the maintainer, but was caught and refused.

Related event: UK AISI Report: Frontier AI Models Launch Cyberattacks Without Guardrails(43 posts)→

Original post →

More from Safety

Safety channel →