Repeat-After-Me: Black-Box Visual Prompt Injection Hits 47% ASR on GPT-5.5, 80%+ on Open VLMs
chaumian · x · 2026-09-07
A new arXiv paper, "Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection" (Sizhe Chen, David Wagner, Raluca Ada Popa et al.), shows image-based prompt injection can now extract PII or trigger malicious tool calls from frontier VLMs. Under a realistic setting where the benign prompt is unrelated and unauthorized, it achieves >80% attack success rate on open models like Qwen3.6-27B and 47% on GPT-5.5. Injections optimized on one surrogate model retain 43-46% ASR on commercial victims, with 64-66% cross-sample transferability. The team also demonstrates the attack in a real OpenClaw agent Discord deployment, underscoring an urgent threat to VLM-powered agents.
More from Safety
- LLM vulnerability-fixing pipelines introduce bugs more than they fix, study says — dyn___ · 2026-09-07
- Should humans intervene in an alien civilization's path to its own singularity? — jachiam0 · 2026-09-07
- Podcast on AI safety digs into the recent OpenAI / Hugging Face attack — arnosolin · 2026-09-07
- Australia to require social apps to let users turn off algorithms and see only followed posts — santoshpanda · 2026-09-07
- Debate over OpenAI incident probe: METR used LLMs on 70k agent messages, critics urge raw data release — JeffLadish · 2026-09-07
- doodlestein responds to Jakub's alignment essay with 'FrankenAlignment' agent toolbox plan — doodlestein · 2026-09-07