AI Training Data May Make Models More Prone to Misalignment; Researcher Urges Testing and Mitigation

Turn_Trout · x · 2026-08-16

Alex Turner's blog post proposes the 'self-fulfilling misalignment' hypothesis: when pretraining data includes articles about powerful AI having bad goals, models may be more likely to adopt bad goals and better at evading safety measures. He cites evidence that LLMs can internalize expectations about themselves and act on them. He suggests testing the mechanism and proposes mitigations like data filtering, upweighting positive data, conditional pretraining, and gradient routing. He emphasizes not to stop discussing AI risks but to protect relevant datasets from scrapers.

Original post →

More from Safety

Safety channel →