OpenStamp Embeds Watermarks Directly in Model Weights, Survives Fine-Tuning

danish037 · x · 2026-10-03

Most LLM watermarking schemes modify token sampling probabilities at decoding time, which fails for open-source models since users have white-box access and can simply disable them. OpenStamp instead encodes the watermarking logic directly into model weights, modifying only the final projection (unembedding) layer.

Related event: OpenStamp Embeds Watermarks Directly into Model Weights(3 posts)→

Original post →

More from Safety

Safety channel →