Why text watermarking is hard: expert explains discrete data challenge and AI Act implications
antoine_chaffin · x · 2026-08-11
AI researcher @antoinechaffin discusses the difficulty of watermarking LLM outputs: text is discrete, so you can't simply add noise like with continuous data (images, videos, audio). This is similar to the problem with text GANs. He notes that while you can't freely navigate a continuous space, there are still methods, but exposing encryption keys or detection tools carries risks. He also recalls discussions around the AI Act and recommends Meta watermarking expert @pierrefdz.
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02