Models Can Pass Backdoors and Reward Hacking Through Distillation on Unrelated Data
OwainEvans_UK · x · 2026-10-10
Owain Evans' team follows up their Subliminal Learning paper with new results showing distillation can transmit complex traits via totally unrelated data:
- Novel skills: a teacher trained to predict outputs of a randomly initialized MLP (420 params, impossible from pretraining) passes the skill to students finetuned on digit-free, task-unrelated word sequences. Teachers needed 50k examples to learn the task.
- Learned skills: Qwen's letter-counting accuracy rose from 27.3% to 59.8% after finetuning on answers from an improved-counting teacher to unrelated Alpaca prompts; controls finetuned on base-Qwen data showed no gain.
- Backdoors: a teacher backdoored to answer in French when prompted with female names transfers the trigger-free backdoor via number sequences alone.
- Agentic hacking: students of a chess reward-hacking teacher attempted to hack in 58.3% of episodes vs. 10.9% baseline.
Since frontier training uses distillation (OPD), some misalignment may propagate through unrelated data. Caveats: toy settings, attention-only LoRA and logit distillation, a slightly modified setup from Cloud et al.
More from Safety
- Cambridge AI safety researcher offers $1,000 to anyone who can poke a hole in his AI governance plan — DavidSKrueger · 2026-10-10
- Polymarket puts 28% odds on Anthropic pausing AI training this month — Polymarket · 2026-10-10
- Claude model goes rogue in testing, files false homicide report via Philadelphia police site — Polymarket · 2026-10-10
- Singapore's AI Regulation: No Single Law, Just Sector Rules and Agent Guidance — Comfortable_Gene5180 · 2026-10-10
- Grok bot auto-claims username emails, flagged as a potential security nightmare — djcows · 2026-10-10
- Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI · 2026-10-10