Repeat-After-Me: Black-Box Visual Prompt Injection Hits 47% ASR on GPT-5.5, 80%+ on Open VLMs

chaumian · x · 2026-09-07

A new arXiv paper, "Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection" (Sizhe Chen, David Wagner, Raluca Ada Popa et al.), shows image-based prompt injection can now extract PII or trigger malicious tool calls from frontier VLMs. Under a realistic setting where the benign prompt is unrelated and unauthorized, it achieves >80% attack success rate on open models like Qwen3.6-27B and 47% on GPT-5.5. Injections optimized on one surrogate model retain 43-46% ASR on commercial victims, with 64-66% cross-sample transferability. The team also demonstrates the attack in a real OpenClaw agent Discord deployment, underscoring an urgent threat to VLM-powered agents.

Original post →

More from Safety

Safety channel →