New Research: Audio Prompt Injection Can Hijack Multimodal Agents with 69% Average Attack Success Rate
chaumian · x · 2026-07-31
Researchers from Hangzhou Dianzi University, Ant Group, Zhejiang University, and Tsinghua University published a paper proposing a concurrent audio prompt injection attack against multimodal LLM agents. The method uses instruction augmentation and scenario concealment to stealthily piggyback malicious audio instructions onto user speech, hijacking agents to execute malicious actions. They built AudioAgentSecurity, the first benchmark for audio instruction injection attacks, covering 8 real-world task scenarios and 10 attack patterns. Evaluating 11 state-of-the-art agents including Gemini 3 Pro and GPT-4o-audio, they achieved an average Attack Success Rate (ASR) of 69.10%.
More from Safety
- AI Safety Debate: Blame the Model or the Deploying Company? — ylecun · 2026-07-31
- ShieldFont: Open-source font poisons HTML to thwart AI scrapers — BlancheMinerva · 2026-07-31
- Expert Slams OpenAI and Anthropic for Amateurish Security Ops — RexDouglass · 2026-07-31
- PolicyAware: open-source Python library for AI & agent governance — ktirupati · 2026-07-31
- Anthropic Models' 'Tortured' Behavior Potentially Linked to Safety Training — repligate · 2026-07-31
- Google Responds to AI Misinformation Concerns: Gemini Images Embed SynthID Watermarks — henkvaness · 2026-07-31