New paper: LLMs transmit traits via unrelated data, and the effects can be proactively detected

StanfordAILab · x · 2026-09-23

A new paper by Nathan Hu with Chris Potts and Sanmi Koyejo studies subliminal learning in LLMs: models transmit traits (e.g. loving cats) through seemingly unrelated data such as numbers. The team proactively detects these effects as readable prompts, leveraging the surprising ability of models to verbalize learned soft prompts. Koyejo highlights Section 5: the trait is recoverable from the data even when the student model doesn't learn it.

Original post →

More from Safety

Safety channel →