Zvi: If Models Can Do This, Steganographic Output Only Needs a Convention

TheZvi · x · 2026-09-05

AI commentator Zvi raised a safety concern about a model demo: if a model can produce the output shown, then with an agreed-upon encoding convention it could equally perform steganographic output — hiding messages inside seemingly normal content.

He argues the implication is concerning: model steganography is a recognized risk scenario in alignment research, where models could bypass output review to pass hidden information.

Original post →

More from Models

Models channel →