AI watermarking detection relies on original logits and prompts
binarybits · x · 2026-08-16
A technical discussion on the feasibility of AI watermarking detection: Is access to the logits from the original generation required? Without the full original prompt, how does one identify which words to inspect for the watermark? It highlights the difficulty of detection without original logits.
More from Safety
- Polymarket: 69% Chance a US State Enacts Data Center Moratorium by End of 2026 — Polymarket · 2026-08-16
- Strong AI Governance Becomes Key Competitive Advantage Over Raw Compute — noahsolomon · 2026-08-16
- Scaling laws are predictable, but risk is jagged and concentrates unexpectedly — chrisrohlf · 2026-08-16
- Google experimented with text watermarking back in 2011 — yoavgo · 2026-08-16
- Who is accountable when AI agents make bad decisions? — KKevinjad · 2026-08-16
- Anthropic addresses watermarking concerns after false positive reports — deliprao · 2026-08-16