New paper: SALVE detects subliminal trait transmission in LLMs before it strikes

ChrisGPotts · x · 2026-09-18

A new paper studies subliminal learning, where LLMs transmit traits (e.g. loving cats) through seemingly unrelated data (e.g. numbers), opening an unnerving avenue for data-poisoning attacks. The proposed SALVE method exploits models' surprising ability to verbalize learned soft prompts as readable prompts, proactively detecting these hidden effects before they cause harm.

Original post →

More from Safety

Safety channel →