Discussion on AI Watermark Detection: Training Black-box Classifiers vs. Reverse Engineering

rasbt · x · 2026-08-16

Regarding AI watermark detection, Sebastian Raschka suggests a potential approach: collecting millions of watermarked and non-watermarked texts to train a binary black-box classifier that identifies hidden patterns. Another perspective argues that detection remains difficult even with full prompt and logit distributions, as the output resembles non-watermarked text. Successful detection might require a massive corpus with distribution data to reverse engineer the token-dependent random seed and position.

Related event: Why Detecting AI Text Watermarks Is So Hard(6 posts)→

Original post →

More from Safety

Safety channel →