Stanford's SALVE Decodes Hidden Trait Transfer in LLMs
Stanford researchers propose SALVE, a method to detect and translate subliminal learning in LLMs, where traits can be covertly transmitted through seemingly unrelated data, posing new data poisoning risks.
2026-09-18 ~ 2026-09-20 · 2 related posts
- New paper: SALVE detects subliminal trait transmission in LLMs before it strikes — ChrisGPotts · 2026-09-18
- Stanford's SALVE decodes hidden behaviors smuggled in distillation data — burny_tech · 2026-09-20