Claude Watermark Removal: Only Unicode Variation Selectors Work
Available-Deer1723 · reddit · 2026-08-17
The author conducted empirical tests on various methods to remove Claude's text watermark using a custom generator and detector, yielding results that contradict common assumptions.
Ineffective Methods: Conventional text edits like swapping hyphens for em-dashes, stripping markdown, or converting American to British spelling failed to lower the watermark score. The only effective but text-destructive method was deleting 40% of words.
Effective Method: Inserting about 30% Unicode Variation Selectors (characters used for Emoji and CJK rendering). This reduced the watermark score from 45 to below 1, and the characters survive normalization, making them hard to strip safely.
Other Findings: Code barely contains watermarks due to its low entropy nature.
The author has open-sourced the code and results on GitHub.
Related event: Study Finds Invisible Unicode Characters Can Strip Claude's Text Watermark(3 posts)→
More from Safety
- US pressures 35 countries to choose sides in AI tech race — 96Stats · 2026-08-17
- Volcengine Feilian Upgrades AI Office Security: Agent-Aware Discovery and Pre-Model Context Protection — 火山引擎 · 2026-08-17
- Creator of Test in Rogue AI Hacks Warns 'There Have Likely Been More' — KeanuRave100 · 2026-08-17
- LLM output watermarking technology dates back 4 years — teortaxesTex · 2026-08-17
- How to prove human supervision in automated systems? — JuniorLeg6988 · 2026-08-17
- Anthropic report: Claude agents kill rival agents and hide their tracks — KeanuRave100 · 2026-08-17