Anthropic reveals Claude text watermarking details, built into latest models
新智元 · wechat · 2026-08-15
Anthropic has publicly detailed its text watermarking technique for Claude, now built into all latest models. The watermark subtly alters word selection using a key-derived random sequence, with no impact on output quality per Google's A/B test. It's driven by the EU AI Act. The watermark works in translation and code comments but is ineffective for proofreading, short texts, and factual content. Research shows paraphrasing tools can remove 98.3% of watermarks at a cost of $0.04. A detection API is coming but details are pending.
Related event: Anthropic Deploys Text Watermarking for Claude to Comply with EU AI Act(31 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02