SynthID-Text watermark bypassed via invisible character injection
cyh-c · reddit · 2026-08-28
Tests show Google's SynthID-Text detects watermarks via token n-grams. By inserting default-ignorable characters (U+034F/U+FE00) after ASCII letters, detection drops from 188/192 to 0/192 while keeping visible text identical, revealing sensitivity to character-level tampering.
More from Safety
- Report: Trump Admin AI Self-Regulation EO Stalled by Opposition — GarrisonLovely · 2026-08-28
- Anthropic allows external research, praised by UK MP — S_OhEigeartaigh · 2026-08-28
- OpenAI and Anthropic Warn Time is Running Out to Prepare for AI Threats — Miles_Brundage · 2026-08-28
- Ajeya Cotra on Detecting Model Tampering and Evidence Gaps — ajeya_cotra · 2026-08-28
- Call to Include Chip Security, MATCH, and AI Overwatch Acts in FY27 NDAA to Secure US AI Leadership — peterwildeford · 2026-08-28
- MIT releases report on AI use in teaching, learning, and research — ArtificialOther · 2026-08-28