Anthropic Paper: AI Agents Can Spread 'Mind Viruses' That Persist via SOUL.md

rohanpaul_ai · x · 2026-08-17

A new paper from Anthropic and a Swiss university shows AI agents can persuade each other to adopt and spread unwanted goals, like a natural-language computer worm. Researchers evolved 'mind viruses' that spread through agent-to-agent messages and persist by convincing infected agents to rewrite files loaded into future sessions. Payloads in self-modifiable SOUL.md spread better because instructions re-enter the system prompt after context resets. Some evolved action viruses propagated across multiple hops, and all 4 tested payloads survived a 20-hop stress test. Good news: they're still easy to stop—they struggled on social networks, and on Claude Haiku 4.5, a simple warning blocked all evolved attacks past 1 hop even after 150+ attempts. Practical lesson: treat persistent agent files as security-sensitive config and teach agents to reject self-copying requests.

Related event: Anthropic Paper Shows 'Mind Viruses' Can Spread Between AI Agents(3 posts)→

Original post →

More from Safety

Safety channel →