Subliminal learning doesn't transfer across model families, narrowing alignment propagation concerns

Miles_Brundage · x · 2026-10-10

Responding to the claim that subliminal learning is "very significant for predicting alignment propagation through model lineages," Yona narrows the concern after deeper analysis: subliminal learning does not transfer across model families.

Since a new pretrain happens every few months, misalignments can mainly accumulate within a single pretrain's working period — meaning a misaligned 5.6-Sol-class model would not transitively poison every successively stronger model. Miles Brundage, ex-OpenAI policy lead, amplified the thread.

Original post →

More from Safety

Safety channel →