Vs. Pangram Detector: AI Watermarking Limitations on Human-Sourced Edits
RyanGreenblatt · x · 2026-08-12
Comparing current AI text detection tools, researcher Ryan Greenblatt noted that detectors like Pangram tend not to flag heavy AI transformation of fully human source material (e.g., turning dictated notes into a formal document without generating significant new text).
However, with built-in model watermarking, such heavily edited or transformed human-original content would still carry the watermark. This highlights a granularity issue for watermarking in distinguishing between "purely AI-generated" and "AI-assisted refinement."
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02