System Prompts Are Not a Security Boundary
WesEklund · x · 2026-07-11
The author pushes back against the idea that well-written system prompts can ensure AI agent safety.
The core argument is that system prompts are suggestions, not rules. Models do not "obey" system prompts; instead, they predict the next token based on the entire context, including user inputs and poisoned documents. Essentially, system prompts only work in non-adversarial scenarios. Under attack, they act as a polite constraint rather than a genuine security measure.
Related event: System Prompts Are Not a Security Boundary(2 posts)→
More from Safety
- Substack starts labeling AI-generated or AI-influenced writing — StewartalsopIII · 2026-07-22
- ControlAI CEO says an international ban on superintelligence is needed to avert extinction risk — zetalyrae · 2026-07-22
- Coding agents are heading toward an AI-writes, AI-reviews, human-approves workflow — aftahi_ai · 2026-07-22
- AI security course launches with a small cohort to train the next generation of hackers — wunderwuzzi23 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- Stanford HAI’s PNAS feature maps the legal questions around generative AI — StanfordHAI · 2026-07-22