Anthropic explains how Claude's text watermarking works

surprisetalk · hn · 2026-08-15

Anthropic published a blog post detailing the mechanism behind Claude's text watermarking. The technique embeds a verifiable signal by statistically adjusting token sampling probabilities with minimal impact on text quality. The post highlights its effectiveness in identifying AI-generated content while discussing trade-offs regarding output quality and robustness challenges.

Original post →

More from Models

Models channel →