Interactive Demo Explains Anthropic's Watermarking via Probability Bias

VeryWellVersed · x · 2026-08-21

Ian Lurie published an interactive demo explaining the likely mechanism behind Anthropic's watermarking. LLMs select words based on probability lists; watermarking subtly tweaks this list to bias the likelihood of certain words. Detectors analyze the text for this specific pattern shift (a 'green list' of words) to determine if it was generated by Claude. The demo lets users adjust bias intensity and observe changes in word probabilities and detection metrics like z-score.

Original post →

More from Safety

Safety channel →