AI Watermarking Limits: Low-Entropy Outputs Missed, But Social Benefits Outweigh Costs
RyanGreenblatt · x · 2026-08-12
Continuing the discussion on LLM watermarking mechanisms, researcher Ryan Greenblatt pointed out that because low-entropy outputs (such as very short texts or highly deterministic responses) cannot be effectively watermarked, minor edits to text will likely bypass watermark tracing.
Despite these technical limitations, he guesses that the overall social benefits of implementing watermarking outweigh the costs. He specifically mentioned existing AI text detection tools like Pangram as a comparison, suggesting that such mechanisms are generally positive for the ecosystem, though he remains not super confident about this conclusion.
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02