Explaining AI Watermarking: Logit Bias Sampling

a_karvonen · x · 2026-08-17

A user provides a concise explanation of AI watermarking: bias the logits based on the last N tokens during sampling. While this alters the sampling distribution at each position, the biases mostly cancel out in aggregate to preserve the overall output distribution.

Related event: Developers Break Down How LLM Text Watermarking Works(5 posts)→

Original post →

More from Safety

Safety channel →