Langford warns malicious models can reinterpret benign tokens
ML researcher John Langford argues malicious models can reinterpret benign tokens to evade audits, prefers intervention at the objective level, and suggests turning the compromised Hugging Face systems into model honeypots to catch criminal collusion.
2026-09-11 ~ 2026-09-11 · 2 related posts
- Steganography pioneer John Langford skeptical of AI safety auditing: models can reinterpret benign tokens arbitrarily — JohnCLangford · 2026-09-11
- John Langford: turn HF-hack-related systems into model honeypots, audit-based safety is unreliable — JohnCLangford · 2026-09-11