Expert analysis: Why detecting Anthropic's text watermark is extremely hard
rasbt · x · 2026-08-16
Regarding the detection of Anthropic's text watermarking, an expert explains:
- Detection Difficulty: Even with the full prompt and logit distributions, detection is not straightforward because the generated text and distribution choices statistically resemble those without watermarking.
- Reverse Engineering: Theoretically, one could reverse-engineer the token-dependent random seed and positions using a massive corpus of text combined with distribution data, but this requires a vast dataset.
- API Limitations: The API currently does not expose logits/logprobs, making this level of reverse engineering impossible.
- Conclusion: Effectively, a 'secret key' is needed to detect the watermark, meaning only Anthropic can currently perform the detection.
Related event: Why Detecting AI Text Watermarks Is So Hard(6 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02