Encrypted prompt injection: one Copilot model leaked secrets in half the tests

Haunting_Ganache_850 · reddit · 2026-10-07

A Copilot prompt-injection attack exploited a counterintuitive flaw: malicious instructions arrive encrypted, so injection filters only see ciphertext — then Copilot helpfully decrypts them itself, reads local secrets while constructing one of the supplied keys, and sends the secrets out in a network request.

The twist: with the same agent, tools and page, one Copilot model completed the full attack chain in about half the tests while two others refused. With automatic model routing, users may not know which model handled the request.

The takeaway: the model can't be treated as the security boundary — agents that can read files, execute code and make outbound connections need their own controls and monitoring regardless of prompt filters.

Original post →

More from coding & agent

coding & agent channel →