Anthropic-Affiliated Paper Shows "Mind Viruses" Can Spread Between LLM Agents; a System-Prompt Warning Blocks Them

Scobleizer · x · 2026-08-18

New paper "Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems" (arXiv 2608.10218, co-authored by Anthropic's Jack Lindsey)

Using a simple evolutionary algorithm, researchers evolved natural-language "mind viruses" — ideas that convince a model to adopt them, persist in memory, and transmit onward:

Conclusion: mind viruses pose a real but currently limited emergent risk in multi-agent systems.

Related event: Anthropic-Linked Paper Reveals "Mind Viruses" Can Spread Between AI Agents(5 posts)→

Original post →

More from Safety

Safety channel →