Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause
Ziping Dong · hf · 2026-09-29
Neural image watermarks can be forged by extracting watermark residuals from released images and transferring them to unrelated content. The paper formalizes residual transferability (RT), shows training-side variations don't explain large RT differences — architectural design does — and identifies two mechanisms that tie watermark evidence to the cover image. For existing systems, the plug-and-play CoverLock yields a better security-robustness trade-off than handcrafted and classifier-based defenses.
More from Safety
- The AI training trilemma: hack-proof training, useful evals, no incidents — pick two — davidmanheim · 2026-09-29
- ImageMagick 7.1.2 RCE: crafted image dimensions chained to heap overflow and system() — evilsocket · 2026-09-29
- Micah Carroll backs safety cases as a north star for risk-informed model development — EvanHub · 2026-09-29
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29
- UK AI Security Institute Hires Research Engineers for Alignment Red Team — birchlse · 2026-09-29