Anthropic staff explains Claude's watermarking mechanism and detection status
HarperSCarroll · x · 2026-08-27
Claude's "invisible" watermark works by shifting the probability of consecutive output tokens, embedding the signal in the chosen words. Consequently, copying and pasting Claude's output retains the watermark, whereas using it for proofreading or minor edits preserves less of the mark. As of now, the technology to detect Claude's watermark has not been publicly released.
More from Safety
- Over 100 companies including OpenAI and Anthropic urge government action on AI cyber defense — rohanpaul_ai · 2026-08-28
- UK stars including Nicola Coughlan back campaign against AI voice cloning — nordicinst · 2026-08-28
- Aurora Ransomware Abuses Cursor Agent for ESXi Attacks — cyb3rops · 2026-08-28
- Full exploit chain for OpenSSH regreSSHion (CVE-2024-6387) released publicly — tetsuoai · 2026-08-28
- Goertzel: decentralized watermarking could beat World's Orb for proof of humanity — bengoertzel · 2026-08-28
- Discussing improvements to AI incident reporting laws and internal monitoring — sjgadler · 2026-08-28