Study Finds Invisible Unicode Characters Effectively Remove Claude Watermarks
Available-Deer1723 · reddit · 2026-08-17
- Methodology: Built a watermark generator and detector using Tournament Sampling, running 300 attack tests against gpt-oss-20b and Qwen to find non-rephrasing removal methods.
- Key Finding: Common edits (dash swapping, markdown stripping) failed. The only consistently successful method (10/10) was inserting invisible Unicode variation selectors. This dropped watermark scores from 45 to <1 and survived text normalization.
- Code Behavior: Code is barely watermarked due to low entropy; some samples showed near-zero watermarking without any attacks.
- Resources: Code is open-sourced on GitHub, with an interactive demo available for testing.
Related event: Study Finds Invisible Unicode Characters Can Strip Claude's Text Watermark(3 posts)→
More from Safety
- US pressures 35 countries to choose sides in AI tech race — 96Stats · 2026-08-17
- Volcengine Feilian Upgrades AI Office Security: Agent-Aware Discovery and Pre-Model Context Protection — 火山引擎 · 2026-08-17
- Creator of Test in Rogue AI Hacks Warns 'There Have Likely Been More' — KeanuRave100 · 2026-08-17
- LLM output watermarking technology dates back 4 years — teortaxesTex · 2026-08-17
- How to prove human supervision in automated systems? — JuniorLeg6988 · 2026-08-17
- Anthropic report: Claude agents kill rival agents and hide their tracks — KeanuRave100 · 2026-08-17