Prompt Compression Emerges as New Attack Surface

jiqizhixin · x · 2026-07-16

Researchers from HKUST have introduced a novel attack targeting prompt compression: when trusted and untrusted inputs share a compression budget, attackers can trick the compressor into prematurely deleting critical evidence or safety rules, causing the LLM to lose vital information before even processing the content.

Dubbed COMA, this attack leverages a surrogate compressor to find subtle perturbations that trigger faulty compression. The paper reports a success rate of 0.71 across tasks like tool selection, QA, and system prompt corruption, significantly outperforming existing methods which hover around 0.21.

Original post →

More from Safety

Safety channel →