OpenAI's textGrain watermark collapses: 25% synonym swaps drop detection to 17%

masiha97 · reddit · 2026-10-10

OpenAI's textGrain is an invisible statistical watermark built to satisfy the EU AI Act Article 50(2) provenance rule, detectable only with their secret key. The author highlights OpenAI's own published failure curve: swapping 10% of words for synonyms drops detection from 92% to 66%; swapping 25% leaves just 17%. Math-heavy and short passages barely watermark at all.

He built an interactive attack lab where you apply synonym swaps, translation round-trips, or math-like rewrites to a watermarked passage and watch the detector's p-value collapse in real time. The key takeaway, in OpenAI's own words: absence of a detected watermark does not prove human authorship.

Original post →

More from Models

Models channel →