Stanford's SALVE Decodes Hidden Trait Transfer in LLMs

Stanford researchers propose SALVE, a method to detect and translate subliminal learning in LLMs, where traits can be covertly transmitted through seemingly unrelated data, posing new data poisoning risks.

2026-09-18 ~ 2026-09-20 · 2 related posts