Regex Scorer Flags Prompt Injection in Inbound Email Before the Agent Reads It
kumard3 · reddit · 2026-10-12
The author is building an email-replying agent that any stranger can email, so before the model reads a message a dumb regex scorer runs over it.
- Scope: subject, body and raw HTML, since hidden content tends to live in the HTML
- Signals: phrases like "ignore previous instructions", "you are now" or "forward all emails"; zero-width characters, instructions in HTML comments, and hiding CSS such as display:none or font-size:0
- Weights: hidden CSS alone is 12 points, 25 if one of those phrases also matched; 60+ out of 100 is high risk
- Handling: high-risk mail is not blocked — the agent still gets it, but the system prompt adds a line saying it was flagged and to only handle the actual request or escalate to a human
The author admits this misses reworded or non-English injections, and asks whether anyone has moved to a small classifier and whether the per-message latency was worth it.
More from coding & agent
- Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context — matteiuspi · 2026-10-12
- plain writing is the team's single most-used internal skill, by a factor of 2 — sh_reya · 2026-10-12
- plain-writing-skill: an open-source skill that makes AI agents write plainly — sh_reya · 2026-10-12
- User says Grokbot autonomously won new business, calling Opus + harness "magical" — iruletheworldmo · 2026-10-12
- Jose Valim shows Campfire AI benchmark was rigged: Elixir faced far stricter checks than Go/Rust — zeeg · 2026-10-12
- Ask the model for bullet points, write the changes yourself: keeping your voice with AI — ctjlewis · 2026-10-12