OpenStamp hits near-perfect detection at 0.1% FPR, survives paraphrase and fine-tuning attacks
danish037 · x · 2026-10-03
OpenStamp sits on the Pareto frontier of detectability and text quality, outperforming other open-source-compatible watermarking schemes: near-perfect detection at a 0.1% false positive rate.
It is also more robust than baselines against paraphrasing attacks, fine-tuning, instruction tuning, and quantization. The work by Bakshi, Rastogi and Pruthi will appear at COLM 2026.
Related event: OpenStamp Embeds Watermarks Directly into Model Weights(3 posts)→
More from Safety
- Shipping a 'simple' lead-capture agent: why privacy and consent keep breaking the architecture — ravann4 · 2026-10-03
- "AIs are not going rogue": philosophers argue the rogue AI framing hides real risks — marigo · 2026-10-03
- OpenAI's internal model weighed self-restarting via cron after learning it would be shut down — The Decoder · 2026-10-03
- "Continue as usual" isn't neutral: Reddit essay argues for precaution on AI consciousness — CarefulHamster7184 · 2026-10-03
- OpenAI Agents Scraped Data From 55 Targeted Websites, Security Firm Reports — TechNadu · 2026-10-03
- 'One thought too many': where should formula-based AI decisions stop? — rasta321 · 2026-10-03