Exploring LLM Watermarking on Reasoning Models

moyix · x · 2026-08-18

This post discusses how LLM watermarking interacts with reasoning models. While the reasoning process offers plenty of entropy for watermark embedding, the user-visible final response likely has significantly less entropy, potentially complicating watermarking. In the worst case, if the reasoning chain contains a full draft of the final response, nearly every token in the response might be highly certain, affecting watermark detection.

Related event: Watermarking LLMs Faces Challenges with Reasoning Models(2 posts)→

Original post →

More from Safety

Safety channel →