How LLM Text Watermarking Works: GPTZero CTO Breaks Down the KGW Method
mark_k · x · 2026-08-11
The CTO of GPTZero explains how frontier labs like Anthropic, Google, and OpenAI are building text watermarking. Most efficient methods rely on the KGW algorithm:
- Generation: Uses previously generated tokens and a secret key to create a hash. This hash reweights the probability distribution for the next token (e.g., splitting vocabulary into 'green' and 'red' lists).
- Detection: Analyzes the proportion of 'green list' tokens in the text to verify the watermark.
While computationally cheap, this approach has known vulnerabilities and can potentially be defeated.
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02