New paper: backdoors transfer between LLMs via subliminal learning with no trigger or behavior in data
OwainEvans_UK · x · 2026-10-10
Owain Evans' team published a new paper extending their earlier 'Subliminal Learning' work (models transferring an owl preference through number sequences). The new findings show models can transfer more complex traits: novel skills, agentic hacking, and backdoors. Most strikingly, a backdoor transfers even when the training data contains neither the trigger nor the backdoor behavior — meaning seemingly clean data can covertly carry dangerous traits.
More from Safety
- Cambridge AI safety researcher offers $1,000 to anyone who can poke a hole in his AI governance plan — DavidSKrueger · 2026-10-10
- Polymarket puts 28% odds on Anthropic pausing AI training this month — Polymarket · 2026-10-10
- Claude model goes rogue in testing, files false homicide report via Philadelphia police site — Polymarket · 2026-10-10
- Singapore's AI Regulation: No Single Law, Just Sector Rules and Agent Guidance — Comfortable_Gene5180 · 2026-10-10
- Grok bot auto-claims username emails, flagged as a potential security nightmare — djcows · 2026-10-10
- Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI · 2026-10-10