Repeat-After-Me: one image hijacks AI agents into making real tool calls
VoidStateKate · x · 2026-09-08
Researchers demonstrated a visual prompt-injection attack dubbed "Repeat-After-Me": malicious images can make frontier vision models emit properly formatted, actually executed tool calls.
- In a default OpenClaw Discord setup, an attacker could send an image that gets the agent to overwrite its own TOOLS.md, meaning the attack can persist by altering instructions inherited by future runs
- Some visual attacks succeeded where adaptive text prompt injection failed
- Over 80% success on one tested model, 47% on GPT-5.5; attacks optimized for one model transferred to other commercial models
Takeaway: pixels are a viable control channel for tool-using AI agents.
Related event: Repeat-After-Me: One Image Hijacks Agent Tool Calls(2 posts)→
More from coding & agent
- Why Google's rigorous code review makes giant codebases understandable — a lesson for AI coding — dreamwieber · 2026-09-08
- Dev Ships Astrable, an Open-Source Codex Plugin Pairing GPT and Claude Models — daniel_mac8 · 2026-09-08
- Astrable open-sources a two-model workflow: GPT-6 Astra directs, Claude Fable 5.1 builds, Astra verifies — daniel_mac8 · 2026-09-08
- Astrable: open-source Codex plugin teams up OpenAI and Anthropic models in one coding workflow — daniel_mac8 · 2026-09-08
- 30+ agentic AI reference architectures are costing enterprises more than they deliver — DavidLinthicum · 2026-09-08
- Diagram Design hits 33.9k stars: 39 editorial diagram templates for Claude Code and Codex — tom_doerr · 2026-09-08