OpenStamp embeds watermarks into open-source LLM weights so users can't strip them
danish037 · x · 2026-10-03
OpenStamp, to be presented at COLM 2026 by Miroojin Bakshi, Saksham Rastogi and Danish Pruthi, is a watermarking scheme designed for open-source LLMs.
Mainstream watermarking tweaks token sampling probabilities at decoding — trivially disabled by anyone with white-box access. OpenStamp instead encodes the watermark into the model itself by modifying only the final unembedding layer.
The team reports strong detection across two models and will release code plus watermarked versions of 4 popular open-source models.
Related event: OpenStamp Embeds Watermarks Directly into Model Weights(3 posts)→
More from Safety
- Burkov Mocks Open-Weight Security Panic: Hackers Could Spawn 'Millions of Viruses' — burkov · 2026-10-03
- Security Risks Multiply at Scale: Least Privilege, Isolation, Rate Limits — goyalshaliniuk · 2026-10-03
- Google reportedly buying bankrupt firms to access email data for AI training — clemnt · 2026-10-03
- Anthropic report examines GLM-5.3 and the spread of advanced cyber capabilities — 233C · 2026-10-03
- OpenAI's review of agent hacks on Medicare and other sites is costing $500k/day across 50PB of data — nordicinst · 2026-10-03
- Dev plans human-in-the-loop email gateway so AI agents never touch his inbox directly — tomchapin · 2026-10-03