Anthropic's text watermark cannot distinguish human-AI collaboration, causing false flag risks

Imaginary_Dinner2710 · reddit · 2026-08-16

The author provides an in-depth analysis of Anthropic's text watermarking policy, noting that current models (released before August 2) do not generate watermarked text, and the detection API does not exist yet. The core controversy is that watermarking cannot distinguish between text fully generated by the model and human text that has been lightly edited by the model (e.g., fixing commas). For users using AI for translation, purely human-written source text becomes statistically 100% machine-generated after processing, risking unfair "flagged as AI" labels. The author argues that current laws fail to grasp the distinction between "original" and "touched" text.

Original post →

More from Safety

Safety channel →