Claude now watermarks all output text: user reproduces SynthID and open-sources a removal workflow
Imaginary_Dinner2710 · reddit · 2026-09-07
A Reddit user found that Claude (Fable 5.1) now embeds a statistical watermark in all generated text, including translations, based on Google DeepMind's SynthID method — an article written by the author and merely translated via Claude comes back watermarked. Meanwhile, the promised detector is only in private preview for select organizations, so regular users can't check their own text before publishing.
Reproducing SynthID with a self-made test key, the author found:
- A full rewrite removes the watermark but introduces factual errors
- Back-translation (even via Chinese) does not destroy the mark
- Simply asking another model (e.g. Qwen) to rewrite leaves substantial signal intact
The working approach: multiple iterative runs, each time highlighting unchanged 5-gram sequences to guide rephrasing, then fact-checking and fixing errors. The author open-sourced the setup and published a free demo, while cautioning it can't guarantee removal of production watermarks. They also note the mark can't distinguish AI-generated from AI-edited text, disproportionately affecting students, scientists, SEO specialists and non-native English speakers.
More from Safety
- Why AI Control grew: it is highly lab-incentive and lab-story compatible — jankulveit · 2026-09-07
- Seth Lazar: models should be trained to check power, not act as toadies — sebkrier · 2026-09-07
- After the Hugging Face breach, a Reddit user asks why alignment gets so little discussion — ZebraCool · 2026-09-07
- LLM-assisted attacks hit Bitcoin harder, with ~$450M in crypto losses — RSync25 · 2026-09-07
- Fencio GA-launches Shark, an automated red-teaming tool for AI agents — OneSafe8149 · 2026-09-07
- Schmidhuber fires back at OpenAI: concrete RSI algorithms have existed for nearly 4 decades — examachine · 2026-09-07