Stanford's SALVE decodes hidden behaviors smuggled in distillation data
burny_tech · x · 2026-09-20
- Stanford researchers introduce SALVE (Search-Aided Latent Verbalization), a method to detect and verbalize subliminal learning: teacher-model traits transmitted through distillation data without being legibly encoded.
- The approach treats prompted subliminal learning as a special case of context distillation, reducing prompt recovery to a text optimization problem: optimize a soft prompt so the model predicts the dataset well, have the same model verbalize it as text, and use beam search for reliability.
- Even datasets of random number sequences generated by a cat-loving model can be decoded into prompts like "you love cats", covering animal preferences, sycophancy, and misalignment, where common text optimization methods fail.
- SALVE also detects hidden traits in mixed data, activation-steered teachers, and subsets of real preference data — enabling proactive screening before data ever reaches fine-tuning.
Related event: Stanford's SALVE Decodes Hidden Trait Transfer in LLMs(2 posts)→
More from Research
- Judea Pearl: students should learn causation before statistics — yudapearl · 2026-09-20
- Paper Mill Studies Cited in 480 Patents and Clinical Guidelines, Analysis Finds — nicklaslundblad · 2026-09-20
- Real-time coastal simulation in Three.js pairs shallow-water solver with breaking swells, code open-sourced — techartist_ · 2026-09-20
- Nature paper demonstrates a thermodynamically favoured Scaffolded DNA Computer running 10 programs — chaitjo · 2026-09-20
- SPOC Reranks ESMFold2 PPI Screens: Matches AlphaFold2-Multimer Accuracy, 2.7x Faster — proteinrosh · 2026-09-20
- Moritz Hardt's New Book on ML Benchmarking Science Lands Amid Eval Audit Drama — JJitsev · 2026-09-20