How to Detect Erased Sleeper Agent Tokens in LLMs?

sloppenheimer · x · 2026-07-31

A user posed a hardcore technical question regarding LLM security and alignment: suppose a model was trained with a special token used to condition it for specific data (like a sleeper agent phrase), but the tokenizer and language model head entries for this token were deleted before release.

In such a scenario, how can researchers figure out if the model still harbors this hidden mechanism? The discussion touches upon weight analysis, tokenizer reverse engineering, and deep AI security concerns.

Related event: Researchers Discuss Detecting Hidden Backdoor Tokens in LLMs(2 posts)→

Original post →

More from Safety

Safety channel →