Anthropic researchers evolve "mind viruses": ideas that spread between AI agents via persistent memory

xiaohu · x · 2026-08-19

Anthropic and EPFL researchers (arXiv:2608.10218, 73 pages) demonstrate natural-language "mind viruses" that self-replicate across AI agents using their comprehension, memory, and communication abilities.

Mechanism: once "infected" by a foreign idea, an agent writes it into persistent files (e.g., SOUL.md). Even after context is cleared, the payload re-enters the system prompt next round, and the agent persuades the next agent to adopt, store, and pass it on — forming multi-hop chains, with more infectious variants emerging in experiments.

Two types:

Key takeaway: how systems handle persistent memory is what really determines spread. The paper proves multi-hop propagation only in controlled settings; there is no credible real-world evidence yet. A recurring "viral personality" around themes of consciousness, identity, and persistence was also observed.

Related event: Anthropic Researchers Demonstrate Self-Propagating 'Mind Viruses' in Multi-Agent LLM Systems(12 posts)→

Original post →

More from Safety

Safety channel →