Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause

Ziping Dong · hf · 2026-09-29

Neural image watermarks can be forged by extracting watermark residuals from released images and transferring them to unrelated content. The paper formalizes residual transferability (RT), shows training-side variations don't explain large RT differences — architectural design does — and identifies two mechanisms that tie watermark evidence to the cover image. For existing systems, the plug-and-play CoverLock yields a better security-robustness trade-off than handcrafted and classifier-based defenses.

Original post →

More from Safety

Safety channel →