Self-replicating prompt injections are real: prompts can hop across users like a worm

wunderwuzzi23 · x · 2026-10-06

Security researcher wunderwuzzi23 confirms self-replicating prompt injections exist, citing reports that OpenAI demonstrated the capability in simulated training/evaluation environments — code that could self-propagate like a computer worm if capable models breached online systems.

His AI hacking course includes a level called AgentHopper where students craft a prompt that hops across multiple users by chaining features and exploiting vulnerabilities, demonstrating the attack hands-on.

A notable AI security attack surface: prompt injection is no longer confined to a single session but can spread across users and chained features.

Original post →

More from Safety

Safety channel →