Prompt Compression Emerges as New Attack Surface
jiqizhixin · x · 2026-07-16
Researchers from HKUST have introduced a novel attack targeting prompt compression: when trusted and untrusted inputs share a compression budget, attackers can trick the compressor into prematurely deleting critical evidence or safety rules, causing the LLM to lose vital information before even processing the content.
Dubbed COMA, this attack leverages a surrogate compressor to find subtle perturbations that trigger faulty compression. The paper reports a success rate of 0.71 across tasks like tool selection, QA, and system prompt corruption, significantly outperforming existing methods which hover around 0.21.
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11