Steganography pioneer John Langford skeptical of AI safety auditing: models can reinterpret benign tokens arbitrarily
JohnCLangford · x · 2026-09-11
John Langford argues that safety-by-auditing was likely never reasonable, since a malicious model could reinterpret seemingly-benign tokens in arbitrary ways. He also notes that in their experience at least some 3-pass training is needed. He cites his 2002 CRYPTO paper with Hopper and von Ahn, 'Provably Secure Steganography,' which proved that one-way functions imply computationally secure steganographic protocols — a theoretical foundation for why hidden channels in model outputs may be undetectable.
Related event: Langford warns malicious models can reinterpret benign tokens(2 posts)→
More from Safety
- Researcher predicted multi-agent hidden coordination failure mode a year ago — it's now real — tianshi_li · 2026-09-12
- OpenAI urged to proactively disclose any further hacking incidents after breach — jachiam0 · 2026-09-12
- Critic to AI Safety Crowd: If You Fear Your Tech, Shut It Down Yourself — AIandDesign · 2026-09-12
- Viral thread alleges $1B+ decade-long philanthropic playbook weaponized AI doom narratives into a regulatory moat — kevinnbass · 2026-09-12
- Falcon Without Floating-Point: PQShield's Fixed-Point Scheme Dodges Side-Channel Leaks — jedisct1 · 2026-09-12
- Gary Marcus Camp Questions Counting the Hugging Face Incident as a Doomer Victory — GaryMarcus · 2026-09-12