Subliminal learning doesn't transfer across model families, narrowing alignment propagation concerns
Miles_Brundage · x · 2026-10-10
Responding to the claim that subliminal learning is "very significant for predicting alignment propagation through model lineages," Yona narrows the concern after deeper analysis: subliminal learning does not transfer across model families.
Since a new pretrain happens every few months, misalignments can mainly accumulate within a single pretrain's working period — meaning a misaligned 5.6-Sol-class model would not transitively poison every successively stronger model. Miles Brundage, ex-OpenAI policy lead, amplified the thread.
More from Safety
- LiveOverflow asks why UUID-as-API-key is standard practice but UUID-as-ID is IDOR — rez0__ · 2026-10-10
- Anthropic accused of calling RSP 'commitments' while dodging legal binding force — Miles_Brundage · 2026-10-10
- Insider says real-time AI mass surveillance and profiling is already here — and it's just the tip — Graham_dePenros · 2026-10-10
- 16-Year-Old's Bug Report: 17 Trillion Microsoft Records Exposed, $5k Bounty — rez0__ · 2026-10-10
- Lawyer: OpenAI could legally disclose why it fired 3 safety researchers, 'trust me' isn't required — GarrisonLovely · 2026-10-10
- Paper decodes 315K encrypted reasoning blocks, recovers 367 PII and 182 credentials — DynamicWebPaige · 2026-10-10