Text Watermarking Explained: Biasing Logits Based on History

a_karvonen · x · 2026-08-17

The author provides a concise technical explanation of text watermarking: during sampling, bias the logits based on the last N tokens. While this changes the sampling distribution at each position, the biases mostly cancel out in aggregate, mostly preserving the overall output distribution.

Related event: Multiple Technical Authors Break Down How LLM Text Watermarking Works(5 posts)→

Original post →

More from Safety

Safety channel →