How to Detect Hidden Trigger Tokens Deleted from a Model's Tokenizer?
stochasticchasm · x · 2026-07-31
A safety researcher poses a critical question: if a model was trained with a special token to condition it for specific data (like a sleeper agent phrase), but the tokenizer and lm head entries were deleted before release, how can we detect whether the model has such a backdoor?
Related event: Researchers Discuss Detecting Hidden Backdoor Tokens in LLMs(2 posts)→
More from Safety
- Sam Altman on the AI dilemma: trade-offs between loss of control and power centralization — r0ck3t23 · 2026-08-24
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24