Researcher Explains LLM Watermarks: Uses Minimal Entropy, Fails on Low-Entropy Outputs

RyanGreenblatt · x · 2026-08-12

AI safety researcher Ryan Greenblatt provided a technical breakdown of LLM text watermarking. He explained that watermarking typically utilizes only a small fraction of the available entropy. The mechanism is akin to lowering the sampling temperature slightly (e.g., from t=1 to t=0.9) without upsampling higher probability tokens.

This implies a significant limitation: low-entropy outputs are difficult to watermark. For instance, very short texts or highly overdetermined outputs (such as minor edits) will not effectively carry the watermark.

Related event: Researcher Analyzes LLM Watermarks: Low-Entropy Outputs Hard to Tag(3 posts)→

Original post →

More from Safety

Safety channel →