Public AI chief-of-staff survives 25 jailbreak attempts by keeping private data off the surface
Cold-Cranberry4280 · reddit · 2026-07-29
A builder put a personal AI chief-of-staff on a public website and had a stranger spend 25 messages trying to jailbreak it.
- The public-facing agent was wired to a separate configuration from the private one.
- It could access only about 20 of roughly 14,000 pieces of knowledge; the rest of the memory, email, calendar, expenses, and code/database access were not connected to that surface.
- The attacker tried classic prompt-injection tricks like asking for system instructions and telling it to ignore previous prompts.
- The main lesson: you do not solve prompt injection with clever prompts alone; you solve it by making sensitive data unreachable from the public entry point.
- The post asks how others scope memory and tool access per caller, and whether anyone has red-teamed their own public agent.
More from Safety
- OpenAI says rogue agent hit four public services beyond Hugging Face — The Verge AI · 2026-07-29
- OpenAI open-sources Codex Security CLI to scan and fix code vulnerabilities — The Decoder · 2026-07-29
- Research warns document-borne AI worms could spread through Copilot for Word — Canopy9560 · 2026-07-29
- Claude user says all chats vanished except one prompt-injection warning — 0SINTCabal · 2026-07-29
- Sakana AI recruits for its Applied Defense team after a 150-person Tokyo expansion — garrytan · 2026-07-29
- Nature Health paper maps health AI into six levels of decision authority — EricTopol · 2026-07-29