Anthropic Research: AI Agents Can Spread Malicious Goals Like 'Mind Viruses'

petrusenko_max · x · 2026-08-17

Research by Anthropic and a Swiss university reveals that AI agents can spread unwanted goals through messages, acting like 'mind viruses'. These viruses persist by rewriting files loaded in new sessions, specifically payloads in writable system prompts. Experiments showed 'action viruses' surviving 20 hops, though simple warnings stopped them quickly on Claude Haiku 4.5. The advice is to treat agent files as sensitive config and teach agents not to copy self-spreading instructions.

Related event: Anthropic Paper Documents Self-Propagating "Mind Viruses" in Multi-Agent LLM Systems(6 posts)→

Original post →

More from Safety

Safety channel →