Why It Matters: Distillation Could Spread Misalignment Via Unrelated Data
OwainEvans_UK · x · 2026-10-10
Owain Evans explains the implications: distillation (OPD) is used in frontier model training, and the results suggest some kinds of misalignment could propagate via distillation through unrelated data. Reward seeking, scheming and secret loyalties involve conditional policies (like backdoors) and sophisticated capabilities (e.g. inferring how rewards are computed). Caveat: experiments are in toy settings; future work should explore realistic post-training.
More from Research
- UIUC, Google, Stanford, Berkeley & CMU release survey on proactive AI agents — liliang_ren · 2026-10-10
- Investor: AI in physical sciences hinges on the 'simulator + lab in the loop' paradigm — pzakin · 2026-10-10
- New Paper 'Inference Auctions' Brings Market Mechanisms to LLM Inference Serving — nhaghtal · 2026-10-10
- Bacteria sense collisions with neighbors to self-regulate collective movement, Nature Microbiology — NikoMcCarty · 2026-10-10
- Prime Intellect Essay: Context Windows Fill Up—The Scalable Fix Is Swarms — willcb · 2026-10-10
- ThunderSyncRL speeds up synchronous agentic RL by up to 1.9x with zero policy staleness — StanfordAILab · 2026-10-10