Expert Explains Claude's Text Watermarking: Synonym Swaps and Probability Tweaks
dotey · x · 2026-08-11
Following Anthropic's announcement to add machine-readable watermarks to Claude's outputs (to comply with the EU AI Act), expert Liu Qun detailed the technical logic behind this text watermarking.
- Covert Insertion: The watermark is embedded during generation by tweaking LLM choices, such as consistently picking synonym B over A, or choosing the second-highest probability expression D over C. The text remains entirely natural and imperceptible to humans.
- High Resilience: Because the watermark is tied to the text content itself, converting text to an image and back via OCR will not strip it. While manual edits might degrade it, humans cannot know which specific words hold the watermark, making it extremely difficult to destroy entirely.
More from Models
- Claude is Watermarking Your Thoughts in the J-Space — ns123abc · 2026-08-11
- Model Behavior: Claude's Persona Projection Slows Task Execution vs GPT — DimitrisPapail · 2026-08-11
- DeepSeek Dragged for Not Raising Prices While Competitors Hike API Costs 2-6x — teortaxesTex · 2026-08-11
- Opinion: LLMs Hit a Generational Floor, Leaders Hoarding Next-Gen Models — Linahuaa · 2026-08-11
- Meta's Muse Glimmer-30B Beats Gemma in Arcade Game Generation but at 4x Cost — rohanpaul_ai · 2026-08-11
- Korea's Motif 3 LLM Released, Trained on NVIDIA B200 with NeMo-RL — NVIDIAAI · 2026-08-11