Anthropic details Claude's text watermarking: using secret keys to steer word choice for EU compliance
heypearlai · x · 2026-08-15
Anthropic has released a blog post detailing the text watermarking technology slated for future Claude models, designed to comply with the EU AI Act. The method involves using a secret key to steer the model's choice when multiple candidate words have similar probabilities (e.g., preferring "overcast" over "grey"), thereby embedding an invisible marker without changing the text's visible content. Anthropic claims this approach has no practical impact on quality, adds no cost, and contains no traceable metadata. The post highlights a user's critique that this technique effectively manipulates the model's inherent randomness.
Related event: Anthropic Deploys Text Watermarking for Claude to Comply with EU AI Act(31 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02