New paper: SALVE detects subliminal trait transmission in LLMs before it strikes
ChrisGPotts · x · 2026-09-18
A new paper studies subliminal learning, where LLMs transmit traits (e.g. loving cats) through seemingly unrelated data (e.g. numbers), opening an unnerving avenue for data-poisoning attacks. The proposed SALVE method exploits models' surprising ability to verbalize learned soft prompts as readable prompts, proactively detecting these hidden effects before they cause harm.
More from Safety
- AI slowdown debate crashes Dreamforce as OpenAI, Anthropic and Nvidia CEOs clash over pacing — nordicinst · 2026-09-18
- Polling shows supermajority support for AI regulation, contra X sentiment — GaryMarcus · 2026-09-18
- Unredacted filings: Microsoft scientist called LLM training "an astonishing theft of unprecedented proportions" — GarrisonLovely · 2026-09-18
- Sen. Hawley flatly rejects antitrust waiver sought by Anthropic and other AI companies — AlexTensor · 2026-09-18
- Research finds AI watermarking like SynthID-Text shifts LLM behavior and can weaken safety guardrails — Ars Technica AI · 2026-09-18
- Rob Leclerc: model unmonitorability is a lab choice, not an inevitability — robleclerc · 2026-09-18