Attack your own AI agent before shipping: multi-turn attacks with harmless-looking failures
iayanpahwa · reddit · 2026-09-08
A practical security walkthrough shows how to red-team your own tool-using AI agent before shipping.
- Demoed against an intentionally weak agent, broken via multi-turn attacks that gradually push it off rails
- Key insight: some failures look completely harmless turn by turn — only the full conversation reveals the agent was manipulated
- Argues adversarial testing should be a standard part of agent release; full steps in the linked blog
More from coding & agent
- State Machines launches agent infrastructure that recreates enterprise apps as stateful API replicas — iamfakhrealam · 2026-09-09
- The hardest part of enterprise agents is knowing they won't break real systems — SimplyAnnisa · 2026-09-09
- Greg Kamradt: AGI's 'general' means doing what it wasn't trained for, via self-built tools — GregKamradt · 2026-09-09
- Meta's Muse Spark 1.3 passes DeepSeek V4 Flash on opencode with 23T tokens used — jack_w_rae · 2026-09-09
- agent-browser adds 60fps recording, targeting review/QA as the new AI coding bottleneck — moeinteractive · 2026-09-09
- Two Codex sessions talk to each other: chief of staff meets supervisor agent — Dimillian · 2026-09-09