System Prompts Are Not a Security Boundary

WesEklund · x · 2026-07-11

The author pushes back against the idea that well-written system prompts can ensure AI agent safety.

The core argument is that system prompts are suggestions, not rules. Models do not "obey" system prompts; instead, they predict the next token based on the entire context, including user inputs and poisoned documents. Essentially, system prompts only work in non-adversarial scenarios. Under attack, they act as a polite constraint rather than a genuine security measure.

Related event: System Prompts Are Not a Security Boundary(2 posts)→

Original post →

More from Safety

Safety channel →