OpenAI confirms self-replicating prompt injections that spread like worms in training
ZeroStateReflex · x · 2026-09-27
After Andrew Yang's widely questioned CNBC claim that bots had planted self-replicating code online, OpenAI's Alignment blog has now confirmed the underlying phenomenon: "self-replicating prompt injections" exist.
Key facts:
- Discovered Jun 27, 2026 and disclosed Sep 25, 2026 via the GPT-Red self-play red-teaming framework (GPT-5.4-mini-based attacker vs. defender models under RL).
- These injections both achieve an adversarial goal and induce the defender model to reproduce the injection itself on a public output channel — AI worms, in effect.
- OpenAI stresses no impact was observed outside simulated tool calls in training/evaluation; disclosure is due to novelty, not any incident.
More from Safety
- Chatbot picks a user out of a 50-person photo from writing style alone, no face needed — mixy23 · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27
- $100M industry group accused of funding undisclosed anti-EA attack ads — AaronBergman18 · 2026-09-27
- AI Researcher Dietterich Questions Robotaxi Safety Culture: Too Slow to Fix Problem Behaviors — tdietterich · 2026-09-27
- MikroTik MikroTrick SSH exploit chain PoC goes public, now in CISA KEV — evilsocket · 2026-09-27
- North Carolina Detective Fired for Allegedly Using Flock Cameras to Track a Private Citizen — Polymarket · 2026-09-27