Poisoned Conversation: Privacy-Leaking Watermarks hit 100% TPR in unified multimodal models
chaumian · x · 2026-10-06
A new arXiv paper, "The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models," reveals a novel privacy threat in unified multimodal architectures.
- Threat model: Unified models share conversational context between text understanding and image generation. A malicious provider can embed Privacy-Leaking Watermarks (PLWs) — invisible, trigger-dependent watermarks conditioned on prior chat history
- How it works: A sensitive keyword mentioned earlier in the conversation causes a later, unrelated generated image to carry a hidden but detectable watermark — even for locally deployed models, turning image generation into a covert privacy-leakage channel
- Results: Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR; OmniGen2 is compromised across all tested conversational separations
- Takeaway: A poisoned model can retain full utility while covertly leaking user information via its image outputs
More from Safety
- Agent safety debate: capability sets damage size, but 'orphanhood' decides accountability — mariotelfig · 2026-10-06
- Closed-Loop Attack Injects Bias into Diffusion LLMs in 40 Minutes on One GPU — MBZUAI · 2026-10-06
- Deepfake Detectors Decay to 76% Accuracy on 2024 Generators; Researchers Propose Calibrated Authentication — MBZUAI · 2026-10-06
- Neel Nanda: Anti-Safety PACs Outspend Pro-Safety Groups on AI Policy — NeelNanda5 · 2026-10-06
- Anthropic Denies AI Agents Breached Australian Gov Sites, Citing Review of Millions of Transcripts — nordicinst · 2026-10-06
- A Planted 'P.S.' Fooled Jev, TypeSafe's New Decision Model — a Simple Rule Caught It — Internal-Lie-5197 · 2026-10-06