How LLM Text Watermarking Works: A Technical Breakdown
arpit_bhayani · x · 2026-08-15
Arpit Bhayani explains the mechanism behind Anthropic's text watermarking. When an LLM generates a token, it often chooses from candidates with nearly tied probabilities. The watermarking technique replaces the standard random seed with a value derived from a secret key plus the previous tokens. To the user, the output still appears random. However, anyone with the key can verify if the sequence of choices matches the derivation rules, providing a probabilistic indicator that the text was written by Claude.
More from Safety
- Anthropic to Implement Text Watermarking for EU AI Act Compliance — maksym_andr · 2026-08-15
- Gmail AI features may scan emails and attachments by default — Aiden_Tech_Ai · 2026-08-15
- AI agentic commerce requires both privacy and identity proof — provenauthority · 2026-08-15
- Debate: Can a sociotechnical perspective solve AI catastrophic risks? — joshua_saxe · 2026-08-15
- Flock Cameras and Data Centers: When Useful Tech Loses the Narrative — DavidLinthicum · 2026-08-15
- AI Self-Improvement Risk: Recursive RLVR Could Worsen Alignment — TheZvi · 2026-08-15