Repeat-After-Me: One Image Hijacks Agent Tool Calls

An arXiv paper demonstrates Repeat-After-Me, a black-box adaptive visual prompt injection attack where a single malicious image induces models to execute real tool calls, achieving a 47% secret-leak rate on GPT-5.5.

2026-09-07 ~ 2026-09-08 · 2 related posts