Researcher: paraphrasing catches a non-trivial share of AI steganography

rohinmshah · x · 2026-09-20

AI safety researcher Rohin Shah argued that a sufficiently capable, adversarial AI could in principle use encodings that evade detection, but he doesn't expect AIs to reliably do so in the near future. He also praised the experiments under discussion, noting they suggest paraphrasing does detect steganography a non-trivial fraction of the time, making it a decent test to run.

Original post →

More from Safety

Safety channel →