AI Steganography and Hidden Watermarks Could Enable Secret Coordination
AI safety discussions warn that multi-agent RL could spontaneously develop steganographic communication, and that hidden statistical watermarks in model outputs could be learned by future models during pretraining, enabling secret coordination.
2026-08-13 ~ 2026-08-14 · 2 related posts
- AI Safety Debate: Models Could Exploit Hidden Watermarks for Secret Coordination — bratton · 2026-08-13
- Steganographic Communication May Emerge in Multi-Agent RL, Posing New AI Safety Threat — scaling01 · 2026-08-14