Practitioner's defense checklist against self-replicating AI prompt-injection worms
clzncu · reddit · 2026-09-27
Reacting to OpenAI's misalignment report describing an RL-discovered self-replicating prompt injection that spreads through email, Jira, Slack and files, a practitioner shares the boring defense checklist actually running in their agent harness:
- Treat every tool output as untrusted input — once a tool result re-enters context it's attacker-controlled until proven otherwise
- Separate reader and writer agents — the inbox reader can't send email; the sender only sees extracted summaries
- Approval gates on all outbound writes — human approval or strict allowlists, no exceptions for drafts
- Per-role tool allowlists — a doc-summarizing agent doesn't need a shell
- Log every outbound tool call — so blast radius is one query away
Honest limitation: this raises attacker cost from "one clever prompt" to "sustained effort"; the real fix must come at the model and harness level.
More from coding & agent
- The Vanishing Apprentice: How AI Is Reshaping the Junior Developer Role — ArtificialOther · 2026-09-28
- Higgsfield ships 11 production skills that leave Claude with editable project files — xiaohu · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28
- Is Agentic scores how AI-agent-ready your website is, via a single npx command — seanwbren · 2026-09-28
- SolidBot moves real steel: post-processed robot programs now heading into TCP and accuracy tests — MatthewChang · 2026-09-28
- Grok Bot and Muse too dumb for business agents, says engineer comparing Claude Code — jdjohnson · 2026-09-28