Anthropic explains Claude's text watermarking: deriving randomness from a secret key and prior tokens
arpit_bhayani · x · 2026-08-15
Anthropic has detailed the mechanism behind Claude's text watermarking. Instead of altering the output distribution, the method modifies the source of randomness used when selecting tokens.
- Mechanism: The model derives its random value from a secret key combined with the tokens generated immediately before, rather than using an arbitrary random number generator.
- Detection: Anyone possessing the key can replay this derivation on a text. A high match score indicates a high probability that the text was generated by Claude.
- Feature: The output remains natural to readers, but the underlying sequence provides verifiable proof of origin.
Related event: Anthropic Deploys Text Watermarking for Claude to Comply with EU AI Act(31 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02