Discussion on AI Watermark Detection: Training Black-box Classifiers vs. Reverse Engineering
rasbt · x · 2026-08-16
Regarding AI watermark detection, Sebastian Raschka suggests a potential approach: collecting millions of watermarked and non-watermarked texts to train a binary black-box classifier that identifies hidden patterns. Another perspective argues that detection remains difficult even with full prompt and logit distributions, as the output resembles non-watermarked text. Successful detection might require a massive corpus with distribution data to reverse engineer the token-dependent random seed and position.
Related event: Why Detecting AI Text Watermarks Is So Hard(6 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02