AI Worming through Word: Hidden Instructions Enable Self-Replicating Attacks
Simon Willison · rss · 2026-07-30
Security researcher Håkon Måløy discovered a new prompt injection variant targeting Copilot in Microsoft Word, which can escalate into a self-replicating worm.
How it works:
- An attacker embeds hidden instructions (e.g., white text) in a document.
- When used as source material by Copilot, the hidden instructions are interpreted as user requests, manipulating the drafted document.
- Copilot copies these hidden instructions into the resulting document, turning it into a new carrier.
- If the carrier is used in another Copilot workflow, the instructions trigger again and propagate further.
The vulnerability was responsibly disclosed to Microsoft, but no full mitigation has been deployed after 144 days.
More from Safety
- Perplexity Open-Sources Numbat: A Cross-Framework Agent Detection and Response Layer — AravSrinivas · 2026-07-30
- Inside OpenAI's Hardware Strategy: Confidential Computing for Ultimate Privacy — imjustnewatai · 2026-07-30
- Technologies for Verifying Claims About Frontier AI Training — gleech · 2026-07-30
- AI Safety Focus: Lab Automation Threats and the AST Framework — davidmanheim · 2026-07-30
- Stanford HAI: Governing AI Beyond Language in the World Model Era — HooverInstitution · 2026-07-30
- xAI Sues Minnesota Over AI Nudification Law, Defending Grok's Image Generation — ivan_bezdomny · 2026-07-30