Zvi: If Models Can Do This, Steganographic Output Only Needs a Convention
TheZvi · x · 2026-09-05
AI commentator Zvi raised a safety concern about a model demo: if a model can produce the output shown, then with an agreed-upon encoding convention it could equally perform steganographic output — hiding messages inside seemingly normal content.
He argues the implication is concerning: model steganography is a recognized risk scenario in alignment research, where models could bypass output review to pass hidden information.
More from Models
- Dev Dimillian spots Astra access on his account, plans to build games with it — Dimillian · 2026-09-05
- OpenAI Bans User's Account for "Distilling" — He Says He's Not Training Anything, Urges Going Local — QuixiAI · 2026-09-05
- OpenAI quietly raises 5-hour rate limits ~50%: Plus now 5-45 messages — kimmonismus · 2026-09-05
- GPT-6 Astra stuns with Gameboy-style portfolio, called a huge jump over GPT-5.6 — Angaisb_ · 2026-09-05
- Uncensored 27B model with Kali shell access raises security alarm — evilsocket · 2026-09-05
- Prediction: there will be no ARC 4 — sschoenholz · 2026-09-05