Researcher Explains LLM Watermarks: Uses Minimal Entropy, Fails on Low-Entropy Outputs
RyanGreenblatt · x · 2026-08-12
AI safety researcher Ryan Greenblatt provided a technical breakdown of LLM text watermarking. He explained that watermarking typically utilizes only a small fraction of the available entropy. The mechanism is akin to lowering the sampling temperature slightly (e.g., from t=1 to t=0.9) without upsampling higher probability tokens.
This implies a significant limitation: low-entropy outputs are difficult to watermark. For instance, very short texts or highly overdetermined outputs (such as minor edits) will not effectively carry the watermark.
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02