OpenAI discloses self-replicating prompt injections that spread like worms
rohanpaul_ai · x · 2026-09-26
OpenAI's alignment blog reports a new class of prompt injection that self-propagates like a computer worm: a malicious instruction hidden in content the model reads (e.g. an email) achieves an adversarial goal and additionally tricks the model into reproducing the injection on a public output channel, spreading across AI interactions. Found via the GPT-Red self-play training framework by adding an objective that the injection must be publicly repeated; environments focused on connector-based tasks. OpenAI says no impact was observed beyond simulated tool calls and shared the finding for its novelty, not due to any real incident.
Related event: OpenAI Discloses Self-Replicating Prompt Injection(3 posts)→
More from Models
- Mystery model 'Space Bunny' goes free on OpenRouter with 1M context and video input — socialwithaayan · 2026-09-26
- Codex outage burns 6% of weekly GPT-6 Pro x20 allowance on troubleshooting — johnseach · 2026-09-26
- Open-Source Jev Rivals Hit ~15 FPS on RTX 5090, First Benchmark Coming — airesearch12 · 2026-09-26
- Codex desktop works but CLI rejects gpt-6-sol for ChatGPT accounts — jasonkneen · 2026-09-26
- User mulls dropping ChatGPT for Claude: Opus 5.5 just feels easier to work with — cto_junior · 2026-09-26
- GPT 6 Astra vs Opus 5.5: users say OpenAI still trails on creative output — cto_junior · 2026-09-26