New paper shows agents can evade watermarks and AI detectors by stitching base-LM outputs

danish037 · x · 2026-09-29

Researchers bhuwandhingra and danish037 published a new paper identifying a shared weakness of text watermarking and post-hoc detectors like Pangram. A capable agent can sample and stitch together text from a base language model that carries no watermark — even if the agent itself is watermarked — producing output that evades attribution. Post-hoc detectors also struggle with base-model outputs, as prior work has shown. The takeaway: current provenance methods for AI-generated text can be systematically bypassed in agent settings.

Original post →

More from Safety

Safety channel →