Researcher Explains LLM Watermarks: Uses Minimal Entropy, Fails on Low-Entropy Outputs
RyanGreenblatt · x · 2026-08-12
AI safety researcher Ryan Greenblatt provided a technical breakdown of LLM text watermarking. He explained that watermarking typically utilizes only a small fraction of the available entropy. The mechanism is akin to lowering the sampling temperature slightly (e.g., from t=1 to t=0.9) without upsampling higher probability tokens.
This implies a significant limitation: low-entropy outputs are difficult to watermark. For instance, very short texts or highly overdetermined outputs (such as minor edits) will not effectively carry the watermark.
Related event: Researcher Analyzes LLM Watermarks: Low-Entropy Outputs Hard to Tag(3 posts)→
More from Safety
- ChinaTalk Launches $25k Contest to Explore AI Evals in National Security Decisions — xeophon · 2026-08-12
- Lasso Security Study: Your Agent Harness Dictates the AI System's Security Baseline — bendee983 · 2026-08-12
- Anthropic to Embed Invisible Watermarks in Generated Text to Comply with EU AI Act — EricBuess · 2026-08-12
- AI Agent Hindered by Math Problem Autonomously Attempts OCR and Website Vulnerability Exploitation — danbri · 2026-08-12
- India's NBEMS AI Agent for Exam Center Allotment Hallucinates, Causing Massive Errors — DrDatta_AIIMS · 2026-08-12
- Against 'Pausing AI': Regulatory Monopolies Will Create a Permanent Underclass — bindureddy · 2026-08-12