Expert Explains Claude's Text Watermarking: Synonym Swaps and Probability Tweaks
dotey · x · 2026-08-11
Following Anthropic's announcement to add machine-readable watermarks to Claude's outputs (to comply with the EU AI Act), expert Liu Qun detailed the technical logic behind this text watermarking.
- Covert Insertion: The watermark is embedded during generation by tweaking LLM choices, such as consistently picking synonym B over A, or choosing the second-highest probability expression D over C. The text remains entirely natural and imperceptible to humans.
- High Resilience: Because the watermark is tied to the text content itself, converting text to an image and back via OCR will not strip it. While manual edits might degrade it, humans cannot know which specific words hold the watermark, making it extremely difficult to destroy entirely.
More from Models
- Using System One models in Swift: fast, deterministic decisions via Apple Foundation Models — rxwei · 2026-10-03
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- Developer Complains OpenAI's Coding Model Endlessly Scopes Creeps Instead of Finishing Tasks — DavidWells · 2026-10-03
- Sonnet 5 Spotted in Google Antigravity Backend, Which Still Runs Sonnet 4.6 — brandon_galang · 2026-10-03
- Post-training Yandex AliceAI-80B-A3B from scratch: a NaN bug in custom V100 kernels killed one run — jjusko20 · 2026-10-03
- Every's Dev Day chat with Matthew Berman: budget gone the moment he tried Ultrafast — every · 2026-10-03