Could a classifier crack Claude's token-bias watermark without the secret key?

US0RIS · reddit · 2026-09-07

A Reddit discussion proposes attacking Anthropic's Claude watermark, which reportedly works by subtly biasing token selection via a secret key rather than embedding hidden characters.

The idea: collect many Claude and non-Claude responses to identical prompts and train a classifier to distinguish them. Key challenges:

The open question: can enough black-box examples approximate Anthropic's secret-key detector without ever knowing the key? Speculative, no experiments yet.

Original post →

More from Safety

Safety channel →