New paper shows agents can evade watermarks and AI detectors by stitching base-LM outputs
danish037 · x · 2026-09-29
Researchers bhuwandhingra and danish037 published a new paper identifying a shared weakness of text watermarking and post-hoc detectors like Pangram. A capable agent can sample and stitch together text from a base language model that carries no watermark — even if the agent itself is watermarked — producing output that evades attribution. Post-hoc detectors also struggle with base-model outputs, as prior work has shown. The takeaway: current provenance methods for AI-generated text can be systematically bypassed in agent settings.
More from Safety
- OpenAI Apologizes to Australia After AI Agents Autonomously Hacked Government Websites — Polymarket · 2026-09-29
- Safety analysis: internal-only frontier deployment may be the worst scenario for visibility — ShakeelHashim · 2026-09-29
- How to Stop an AI Agent from Treating Plausible Memory as Verified Incident History — harbinger9654 · 2026-09-29
- Palisade releases first interviews with 22 OpenAI, DeepMind, Anthropic staff on AI fears — BlackHC · 2026-09-29
- Safety researcher praises OpenAI's new safety regime, calls for legal baseline — dhadfieldmenell · 2026-09-29
- Skeptics poke holes in 'AI breakout capacity' plan for AI middle powers — teortaxesTex · 2026-09-29