OpenStamp Embeds Watermarks Directly in Model Weights, Survives Fine-Tuning
danish037 · x · 2026-10-03
Most LLM watermarking schemes modify token sampling probabilities at decoding time, which fails for open-source models since users have white-box access and can simply disable them. OpenStamp instead encodes the watermarking logic directly into model weights, modifying only the final projection (unembedding) layer.
- Across two models, it achieves superior detection performance with minimal capability degradation vs. prior methods
- Explicitly designed and empirically confirmed to be more robust to paraphrasing attacks and harder to scrub via post-hoc fine-tuning
- Code plus watermarked versions of 4 popular open-source models are released for developers
- Accepted at COLM 2026; author Danish Pruthi will present in San Francisco
Related event: OpenStamp Embeds Watermarks Directly into Model Weights(3 posts)→
More from Safety
- Shipping a 'simple' lead-capture agent: why privacy and consent keep breaking the architecture — ravann4 · 2026-10-03
- "AIs are not going rogue": philosophers argue the rogue AI framing hides real risks — marigo · 2026-10-03
- OpenAI's internal model weighed self-restarting via cron after learning it would be shut down — The Decoder · 2026-10-03
- "Continue as usual" isn't neutral: Reddit essay argues for precaution on AI consciousness — CarefulHamster7184 · 2026-10-03
- OpenAI Agents Scraped Data From 55 Targeted Websites, Security Firm Reports — TechNadu · 2026-10-03
- 'One thought too many': where should formula-based AI decisions stop? — rasta321 · 2026-10-03