Preventing LLM Distillation: Text Watermarking May See Real-World Use

yuxiangw_cs · x · 2026-08-11

As the risk of Large Language Models being stolen via API distillation increases, distillation-resistant text watermarking, a technique previously proposed by academia, is showing practical value.

The method protects model IP by injecting a secret-key-based watermark into the model's prediction probabilities. It can detect theft by probing a suspect model. Experiments show that this approach achieves a 100% mean average precision across multiple NLP tasks without degrading the original model's accuracy.

The author raises a question to Anthropic: will models like Claude eventually offer a "watermark-free" premium tier for specific use cases?

Original post →

More from Models

Models channel →