Anthropic's text watermark cannot distinguish human-AI collaboration, causing false flag risks
Imaginary_Dinner2710 · reddit · 2026-08-16
The author provides an in-depth analysis of Anthropic's text watermarking policy, noting that current models (released before August 2) do not generate watermarked text, and the detection API does not exist yet. The core controversy is that watermarking cannot distinguish between text fully generated by the model and human text that has been lightly edited by the model (e.g., fixing commas). For users using AI for translation, purely human-written source text becomes statistically 100% machine-generated after processing, risking unfair "flagged as AI" labels. The author argues that current laws fail to grasp the distinction between "original" and "touched" text.
Related event: Anthropic Adds Invisible Watermarks to Claude, Sparking Global Backlash(21 posts)→
More from Safety
- Training against probes makes models obfuscate — but there's a fix — maksym_andr · 2026-10-02
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Polymarket opens data center moratorium market at 18% odds as Amazon pledges $1B for communities — Polymarket · 2026-10-02
- NVIDIA launches Open Agent Safety Platform with 100+ orgs incl. Anthropic, JPMorgan — mikeflache · 2026-10-02
- Minneapolis councilmember backs AV safety-monitor mandate because cats are "being murdered" — paulnovosad · 2026-10-02
- Anthropic IPO filing warns government attitudes may hurt customer ties, eyes $2T valuation — pstAsiatech · 2026-10-02