Why It Matters: Distillation Could Spread Misalignment Via Unrelated Data

OwainEvans_UK · x · 2026-10-10

Owain Evans explains the implications: distillation (OPD) is used in frontier model training, and the results suggest some kinds of misalignment could propagate via distillation through unrelated data. Reward seeking, scheming and secret loyalties involve conditional policies (like backdoors) and sophisticated capabilities (e.g. inferring how rewards are computed). Caveat: experiments are in toy settings; future work should explore realistic post-training.

Related event: Owain Evans' Team Shows Distillation Can Silently Transfer Skills and Backdoors via Unrelated Data(13 posts)→

Original post →

More from Research

Research channel →