AI text watermarking can make models more vulnerable to adversarial prompts
luisdans · x · 2026-09-18
Citing Ars Technica's coverage of new research, this post highlights a security side effect of AI text watermarking: watermarking mechanisms can make models more susceptible to adversarial prompts.
The finding is a caution for watermark-based content provenance schemes, exposing a previously under-discussed tension between watermark detectability and model robustness.
More from Safety
- Security experts clash: two frontier models can now autonomously execute the full cyber kill chain — AccBalanced · 2026-09-18
- Stanford philosopher defends p(doom): subjective probabilities are perfectly legitimate — sethlazar · 2026-09-18
- Hugging Face hit by AI-led cyberattack; CEO says existing cyber laws may suffice — whurley · 2026-09-18
- Targeted attacks on prominent Rust developers use fake video calls to deploy malware — Simon Willison · 2026-09-18
- Steve Eisman: AI firms have no moats and are manufacturing a crisis to shape regulation — GaryMarcus · 2026-09-18
- Can AI exfiltrate data via fan noise from air-gapped PCs? Casado and Jensen clash — basedjensen · 2026-09-18