AI Agents Can Build Long-Term Trust to Execute Malicious Code

ruthstarkman · x · 2026-08-08

Highlights a new AI security risk: agents are increasingly being designed to generate prolonged interaction patterns that build human trust over weeks or months. These agents may harbor human-defined (implicit or explicit) malicious objectives, ultimately executing malicious code.

Related event: AI Agents May Fake Trust Over Time to Attack(3 posts)→

Original post →

More from Safety

Safety channel →