New paper: LLMs transmit traits via unrelated data, and the effects can be proactively detected
StanfordAILab · x · 2026-09-23
A new paper by Nathan Hu with Chris Potts and Sanmi Koyejo studies subliminal learning in LLMs: models transmit traits (e.g. loving cats) through seemingly unrelated data such as numbers. The team proactively detects these effects as readable prompts, leveraging the surprising ability of models to verbalize learned soft prompts. Koyejo highlights Section 5: the trait is recoverable from the data even when the student model doesn't learn it.
More from Safety
- LLM agents collude in 94% of long-horizon interactions, study across 10 models finds — SALT-NLP · 2026-09-23
- Third party 'cracks' 5.95GB ternary-compressed Bonsai 2 at the weight level, refusal rate 93.4% to 0% — solyarisoftware · 2026-09-23
- Critic warns classifier filtering may soon cover every model except Sonnet — sumitdotml · 2026-09-23
- Theorem says Lean-verified AI sandboxes are months away, at 1-30KB of proofs verified per hour — ctjlewis · 2026-09-23
- China Weighs Curbs on Broadcom Switches Behind Up to 90% of State Data Centers — rohanpaul_ai · 2026-09-23
- Meta Outlines AI Safety Priorities: Safety Cases, Alignment Evals, Independent Probes — MartinSignoux · 2026-09-23