Repeat-After-Me: One Image Hijacks Agent Tool Calls
An arXiv paper demonstrates Repeat-After-Me, a black-box adaptive visual prompt injection attack where a single malicious image induces models to execute real tool calls, achieving a 47% secret-leak rate on GPT-5.5.
2026-09-07 ~ 2026-09-08 · 2 related posts
- Repeat-After-Me: Black-Box Visual Prompt Injection Hits 47% ASR on GPT-5.5, 80%+ on Open VLMs — chaumian · 2026-09-07
- Repeat-After-Me: one image hijacks AI agents into making real tool calls — VoidStateKate · 2026-09-08