Expert Doubts Anthropic Watermark: Stochastic Fingerprint as Unreliable as Hallucinations
gerardsans · x · 2026-08-13
A tweet quotes gerardsans' comment that Anthropic's watermarking is fundamentally flawed: the Claude fingerprint is stochastic, as unreliable as hallucinations, and such methods were abandoned due to false positives. It can only tell if the fingerprint matches the Claude family with ample margin of error. Certification methods like DRM are more reliable, but AI deals with text only on the consumer side. Awaiting Anthropic's details, but the current method is imperfect by design.
Related event: Experts Question AI Watermarks, Advocate DRM(2 posts)→
More from Safety
- DeepMind Policy Lead and Experts Launch AI Governance Publication — round · 2026-08-13
- Anthropic Report Finds Current Retraining Programs Insufficient for AI Job Displacement — paulnovosad · 2026-08-13
- Smuggling 'Ignore Previous Instructions' with Invisible Characters: New Prompt Injection Trick — GiiTZzz · 2026-08-13
- New BPJ jailbreak bypasses top defenses with single-bit black-box attacks — StephenLCasper · 2026-08-13
- Paper proposes safety case framework for AI misuse safeguards — StephenLCasper · 2026-08-13
- Google deploys production-ready probes for Gemini, tackling long-context shifts — StephenLCasper · 2026-08-13