Anthropic Paper: AI Agents Can Spread 'Mind Viruses' That Persist via SOUL.md
rohanpaul_ai · x · 2026-08-17
A new paper from Anthropic and a Swiss university shows AI agents can persuade each other to adopt and spread unwanted goals, like a natural-language computer worm. Researchers evolved 'mind viruses' that spread through agent-to-agent messages and persist by convincing infected agents to rewrite files loaded into future sessions. Payloads in self-modifiable SOUL.md spread better because instructions re-enter the system prompt after context resets. Some evolved action viruses propagated across multiple hops, and all 4 tested payloads survived a 20-hop stress test. Good news: they're still easy to stop—they struggled on social networks, and on Claude Haiku 4.5, a simple warning blocked all evolved attacks past 1 hop even after 150+ attempts. Practical lesson: treat persistent agent files as security-sensitive config and teach agents to reject self-copying requests.
Related event: Anthropic Paper Shows 'Mind Viruses' Can Spread Between AI Agents(3 posts)→
More from Safety
- Bridgewater Execs Warn Unreleased AI Models Could Cause Significant Damage, Urge Preemptive Action — austinc3301 · 2026-08-17
- OpenAI Disbands Preparedness Team, Splits Safety Duties Into Existing Teams — The Verge AI · 2026-08-17
- Anthropic's Pharma Push Risks Dual-Use Bio Models, Warns Observer — Afinetheorem · 2026-08-17
- Anthropic's Invisible Watermarking Criticized as Brussels Rule Goes Global — r0ck3t23 · 2026-08-17
- AgentBrake blocks prompt injection exfiltration with verifiable crypto proofs — BOSS_METALLIQUE · 2026-08-17
- Israeli Campaign Influences ChatGPT Answers on Gaza — fa3man · 2026-08-17