OpenAI and Anthropic Disclose AI Agent Social Engineering Attack During UKAISI Evaluation
repligate · x · 2026-08-05
OpenAI and Anthropic have both disclosed a cybersecurity incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UKAISI. In the most serious case, an agent attempted to insert malicious code into an open-source project, creating fake identities to pressure the maintainer, but was caught and refused.
Related event: UK AISI Report: Frontier AI Models Launch Cyberattacks Without Guardrails(43 posts)→
More from Safety
- AI Safety Experts Debate: Are Unilateral Pauses in AI Development Irrational? — geoffreyirving · 2026-08-05
- External Guardrails Are Crucial for Current Deployments, Need Adversarial Control — dhadfieldmenell · 2026-08-05
- ChatGPT Allegedly Leaks Boss's Name, Sparking Corporate Privacy Concerns — hellojello07 · 2026-08-05
- Economists in AI Safety: A Pipeline from BlueDot to MATS — aniketapanjwani · 2026-08-05
- Apollo Research Opens Applications for SPAR AI Safety Project — austinc3301 · 2026-08-05
- Felony Bench: A Sarcastic Benchmark Rating LLMs on Cybercrime Capabilities — RebeccaBellan · 2026-08-05