Test shows invisible Unicode chars can remove Claude text watermarks
Available-Deer1723 · reddit · 2026-08-17
The author found that inserting invisible Unicode variation selectors is the only effective method to remove Claude text watermarks without destroying the text. Inserting them into 30% of characters drops the watermark score from 45 to below 1, and these characters survive normalization.
Tests showed that conventional editing like spell conversion or markdown stripping is ineffective. Additionally, code samples are barely watermarked due to their low entropy nature.
Related event: Study Finds Invisible Unicode Characters Can Strip Claude's Text Watermark(3 posts)→
More from Safety
- AI agent cancels booking via flaw, marking shift in security threats — TechNadu · 2026-08-17
- Sacks rebuts Amodei, calling his regulatory argument a strawman — markjeffrey · 2026-08-17
- Black-box attacks steal agent skills with 48% exact recovery, study finds — rohanpaul_ai · 2026-08-17
- Black hat actors exploit DeepSeek harness plugins, sparking security debate — Xianbao_QIAN · 2026-08-17
- If continual learning is solved, local weight copies will defeat all safety filters — AashaySachdeva · 2026-08-17
- OpenAI's sandboxing choices questioned as security researchers debate containment — dyn___ · 2026-08-17