Expert Doubts Anthropic Watermark: Stochastic Fingerprint as Unreliable as Hallucinations
gerardsans · x · 2026-08-13
A tweet quotes gerardsans' comment that Anthropic's watermarking is fundamentally flawed: the Claude fingerprint is stochastic, as unreliable as hallucinations, and such methods were abandoned due to false positives. It can only tell if the fingerprint matches the Claude family with ample margin of error. Certification methods like DRM are more reliable, but AI deals with text only on the consumer side. Awaiting Anthropic's details, but the current method is imperfect by design.
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02