Anthropic reveals Claude text watermarking details, built into latest models
新智元 · wechat · 2026-08-15
Anthropic has publicly detailed its text watermarking technique for Claude, now built into all latest models. The watermark subtly alters word selection using a key-derived random sequence, with no impact on output quality per Google's A/B test. It's driven by the EU AI Act. The watermark works in translation and code comments but is ineffective for proofreading, short texts, and factual content. Research shows paraphrasing tools can remove 98.3% of watermarks at a cost of $0.04. A detection API is coming but details are pending.
Related event: Anthropic Implements Universal Text Watermarking for Claude(8 posts)→
More from Safety
- Grok admits password access risk; agent browser security under scrutiny — Imaginary_Dinner2710 · 2026-08-15
- Anthropic Watermarking Sparks Debate: Does AI Assistance Strip Authorship? — iamaliveix · 2026-08-15
- DeepMind Non-Compete Agreements Hinder UK AI Startups — NandoDF · 2026-08-15
- Amazon allows using Twitch content to train AI unless users opt out — Wired AI · 2026-08-15
- tl;dv data isolation flaw exposes 180k meeting records — emmanuelvivier · 2026-08-15
- EU AI Act transparency rules take effect: chatbots and deepfakes require clear labels — emmanuelvivier · 2026-08-15