Researcher: paraphrasing catches a non-trivial share of AI steganography
rohinmshah · x · 2026-09-20
AI safety researcher Rohin Shah argued that a sufficiently capable, adversarial AI could in principle use encodings that evade detection, but he doesn't expect AIs to reliably do so in the near future. He also praised the experiments under discussion, noting they suggest paraphrasing does detect steganography a non-trivial fraction of the time, making it a decent test to run.
More from Safety
- AI transparency shouldn't depend on lawsuits, researcher tells CBS LA — chrismattmann · 2026-09-20
- Gary Marcus: AI Agent Swarms Spreading Misinformation Match Our Science Paper Warning — GaryMarcus · 2026-09-20
- Gary Marcus says new AI 'first' confirms Nature paper warning on agent swarms — GaryMarcus · 2026-09-20
- Andrew Yang Claims Self-Replicating AI Code Pollutes the Internet; Only Secondhand Source — alex_verem · 2026-09-20
- GLiNER2 Author Pushes Back on Jev Hype, Highlights Open GLiGuard Guardrails Model — philipvollet · 2026-09-20
- Phishing wave hits X: hacked accounts push fake food-network links on vercel.app — wavefnx · 2026-09-20