Why AI Watermarking May Break Down in Agentic Workflows

yaakg25 · reddit · 2026-09-04

The author argues that AI text watermarking—which depends on statistical context of immediately preceding words—may fail in agentic contexts.

While an agent's first draft can be watermarked, agents typically edit documents via code (e.g., JavaScript/Python patching HTML), so the final PDF's text no longer reflects contiguous model output. The contextual features watermarking relies on get destroyed, causing detection to fail.

Original post →

More from Safety

Safety channel →