Could a classifier crack Claude's token-bias watermark without the secret key?
US0RIS · reddit · 2026-09-07
A Reddit discussion proposes attacking Anthropic's Claude watermark, which reportedly works by subtly biasing token selection via a secret key rather than embedding hidden characters.
The idea: collect many Claude and non-Claude responses to identical prompts and train a classifier to distinguish them. Key challenges:
- The classifier might learn Claude's writing style rather than the watermark
- A stronger design: feed Claude repeated similar prefixes and see if next-token choice patterns are learnable
The open question: can enough black-box examples approximate Anthropic's secret-key detector without ever knowing the key? Speculative, no experiments yet.
More from Safety
- 2,348 alleged Booking.com customer records sold for $40 in Monero, breach unconfirmed — TechNadu · 2026-09-11
- Mozilla CTO calls for major pause on generative AI in schools, warns of losing a generation — Dan_Jeffries1 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- Anthropic Says It Blocked Attempts to Use AI for Bioweapons Development — KoseteBamse · 2026-09-11
- Never hardcode AI API keys: attackers scan app binaries, GitHub and Docker — eyishazyer · 2026-09-11
- Beware hotel Wi-Fi popups: DNS hijacking used to deliver malware — eyishazyer · 2026-09-11