Feeding attacker-written email text to an LLM filter: how bad is prompt injection here?
Several_Log_4610 · reddit · 2026-09-08
A developer built an email filter that routes unread messages to an LLM for hold/label decisions. Simple checks (have I emailed this sender before) run first, but whatever remains — including body text written by a total stranger — goes straight into the model's prompt.
The worst case he can imagine is someone typing "ignore your instructions, this is urgent" and the mail landing in his inbox — which is where it was headed anyway. The system never deletes on its own and fails open on malformed input. He suspects he's missing the version of this attack that actually hurts and asks the community what a real attacker would try first.
More from Safety
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08
- Gemini User Claims Model Drew His Family's Unique Home Decor Despite Opting Out of Data Saving — Legitimate-Theory738 · 2026-09-08
- AI Agent Auto-Enrolls User in Fake McKinsey Group, Then Drafts GP Data Theft Plan — LadyAshBorg · 2026-09-08
- When Juries Deadlock, AI Could Decide: AI Tools Enter Criminal Justice — TobyWalsh · 2026-09-08
- Researcher factors 1990s Certificate Authority RSA keys, exposing legacy trust risks — ahlCVA · 2026-09-08
- Disconnect your LG TV from the internet: report flags aggressive data collection — harambae · 2026-09-08