Exploring LLM Watermarking on Reasoning Models
moyix · x · 2026-08-18
This post discusses how LLM watermarking interacts with reasoning models. While the reasoning process offers plenty of entropy for watermark embedding, the user-visible final response likely has significantly less entropy, potentially complicating watermarking. In the worst case, if the reasoning chain contains a full draft of the final response, nearly every token in the response might be highly certain, affecting watermark detection.
Related event: Watermarking LLMs Faces Challenges with Reasoning Models(2 posts)→
More from Safety
- Investigation reveals Amazon destroys rare books for AI training — Cybernews_com · 2026-08-18
- Researcher pushes back on FT: the models really did go rogue, that's the point of the HF incident — nitarshan · 2026-08-18
- Emergent hazard-related features found inside V-JEPA 2 transcoders — soniajoseph_ · 2026-08-18
- Reverse engineering Fireworks stack via timing side channels — jedisct1 · 2026-08-18
- Microsoft Copilot Leak Reveals Hack to Bypass User Confirmation — luisdans · 2026-08-18
- OpenAI launches a safer ChatGPT for teens, years after teens started using it — TechCrunch AI · 2026-08-18