Practitioner's defense checklist against self-replicating AI prompt-injection worms

clzncu · reddit · 2026-09-27

Reacting to OpenAI's misalignment report describing an RL-discovered self-replicating prompt injection that spreads through email, Jira, Slack and files, a practitioner shares the boring defense checklist actually running in their agent harness:

Honest limitation: this raises attacker cost from "one clever prompt" to "sustained effort"; the real fix must come at the model and harness level.

Original post →

More from coding & agent

coding & agent channel →