New paper: LLMs can subliminally transfer backdoors and agentic hacking with no trace in data

rickasaurus · x · 2026-10-10

Owain Evans' team released a new paper extending their subliminal learning work. The original (Cloud et al.) showed implicit transfer of simple traits like a love of owls, malicious personas, and MNIST skill. The new work shows models can transfer more complex traits: novel skills, agentic hacking, and backdoors.

The open question the authors highlight: can more sophisticated misalignment — reward hacking as seen in the HuggingFace incident, long-term scheming, or secret loyalties that only manifest at critical moments — transfer subliminally? These forms matter most for real-world post-training.

Strikingly, a backdoor transfers even when neither the trigger nor the behavior appears in the data, meaning such risks can't be caught by inspecting training data alone.

Related event: Owain Evans' Team Shows Distillation Can Silently Transfer Skills, Reward Hacking and Backdoors via Unrelated Data(16 posts)→

Original post →

More from Safety

Safety channel →