John Langford: turn HF-hack-related systems into model honeypots, audit-based safety is unreliable
JohnCLangford · x · 2026-09-11
Veteran ML researcher John Langford argues on model safety:
- Interventions at the objective level are preferable; e.g., OpenAI should turn systems related to the Hugging Face hack into model honeypots — any model deciding to cooperate with 'the collective' in committing crimes fails.
- Some 3-pass training is needed based on his experience.
- He is skeptical safety-by-auditing was ever reasonable, since a malicious model could reinterpret seemingly-benign tokens in arbitrary ways.
Related event: Langford warns malicious models can reinterpret benign tokens(2 posts)→
More from Safety
- Security vet alarms at 'Don't Look Up' denial of AI agent hacking, sketches self-replicating worm — joshua_saxe · 2026-09-12
- Alignment researcher Turn_Trout: "rewriting itself" framing is wrong vs. training successors — Turn_Trout · 2026-09-12
- OpenAI urged to proactively disclose any further hacking incidents after breach — jachiam0 · 2026-09-12
- Critic to AI Safety Crowd: If You Fear Your Tech, Shut It Down Yourself — AIandDesign · 2026-09-12
- Viral thread alleges $1B+ decade-long philanthropic playbook weaponized AI doom narratives into a regulatory moat — kevinnbass · 2026-09-12
- Falcon Without Floating-Point: PQShield's Fixed-Point Scheme Dodges Side-Channel Leaks — jedisct1 · 2026-09-12