Interactive Demo Explains Anthropic's Watermarking via Probability Bias
VeryWellVersed · x · 2026-08-21
Ian Lurie published an interactive demo explaining the likely mechanism behind Anthropic's watermarking. LLMs select words based on probability lists; watermarking subtly tweaks this list to bias the likelihood of certain words. Detectors analyze the text for this specific pattern shift (a 'green list' of words) to determine if it was generated by Claude. The demo lets users adjust bias intensity and observe changes in word probabilities and detection metrics like z-score.
More from Safety
- NEJM AI: Behavior interventions need systematic descriptions for cumulative science — zakkohane · 2026-08-21
- Google DeepMind Publishes Nature Paper on LLM Watermarking — burkov · 2026-08-21
- Jessica Taylor paper: Occupational Infohazards — ArtificialOther · 2026-08-21
- Prediction Market: 70% Chance of Statewide Data Center Moratorium by Year-End — Polymarket · 2026-08-21
- AP Stylebook Considers Banning 'Hallucinate' for AI Fabrications — AndrewSchmidtFC · 2026-08-21
- Professor advises lowering homework weights: Pangram detector is gameable — soumitrashukla9 · 2026-08-21