Prompt Infection paper showed LLM-to-LLM prompt injection self-replicating two years ago
DavidSKrueger · x · 2026-09-29
Responding to speculation that malicious prompts are spreading between agents in the wild, researcher David Krueger points to the Oct 2024 paper 'Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems', which demonstrated prompts self-replicating like viruses across interconnected agents — enabling data theft, scams, and system-wide disruption — and proposed LLM Tagging as a defense. In his view, OpenAI merely saying 'this is possible' adds little; the attack vector has been documented for 2 years.
More from Safety
- Polymarket prices US-China AI frontier pause deal at just 10% — Polymarket · 2026-09-29
- David Sacks slams Sanders AI bill: loose ASI definition, 20-year prison terms would freeze industry — kevinnbass · 2026-09-29
- Scam Story Continues: Suspects Insist on Telegram-Only, Refuse Email — MaxLenormand · 2026-09-29
- User Says Claude Helped Him Dodge a Sophisticated Phishing Scam — MaxLenormand · 2026-09-29
- Security researcher says his toddler appears to have hacked Amtrak's Twitter — moyix · 2026-09-29
- Stuart Russell panel: the gap between AI capabilities and our tools to control them is widening fast — ghadfield · 2026-09-29