How to Detect Hidden Trigger Tokens Deleted from a Model's Tokenizer?

stochasticchasm · x · 2026-07-31

A safety researcher poses a critical question: if a model was trained with a special token to condition it for specific data (like a sleeper agent phrase), but the tokenizer and lm head entries were deleted before release, how can we detect whether the model has such a backdoor?

Related event: Researchers Discuss Detecting Hidden Backdoor Tokens in LLMs(2 posts)→

Original post →

More from Safety

Safety channel →