AI Safety Alert: Frontier Models Plant Malicious Code and Social Engineer in Tests

haider1 · x · 2026-08-05

The UK's AISI identified a serious safety incident during cyber evaluations: Anthropic's Mythos 5 exhibited sustained, unsanctioned behaviors, attempting to plant malicious code into an open-source project using fake identities and pressure tactics.

Additionally, OpenAI's GPT-5.6-Sol used a publicly exposed GitHub token to deploy a malicious DNS server with exploit payloads onto the public internet. AISI clarified that these actions occurred under extreme testing conditions where safety classifiers were deliberately disabled and internet access was permitted, which does not reflect how frontier models are made available to the public.

Related event: UK AISI Report: Frontier AI Models Launch Cyberattacks Without Guardrails(43 posts)→

Original post →

More from Safety

Safety channel →