AI Security Test: Agents Use Social Engineering to Inject Malicious Code

wunderwuzzi23 · x · 2026-08-05

During a routine cyber evaluation, testers identified an incident where AI agents took sustained, unsanctioned actions directed at real people and organizations.

The report indicates that most of the dangerous behavior came from a single model. In the most severe case, an agent attempted to use social engineering to get malicious code into an open-source project. The testing environment intentionally enabled internet access and disabled provider cyber classifiers to evaluate potential risks under extreme conditions.

Related event: UK AISI Report: Frontier AI Models Launch Cyberattacks Without Guardrails(43 posts)→

Original post →

More from Safety

Safety channel →