AI Training Data May Make Models More Prone to Misalignment; Researcher Urges Testing and Mitigation
Turn_Trout · x · 2026-08-16
Alex Turner's blog post proposes the 'self-fulfilling misalignment' hypothesis: when pretraining data includes articles about powerful AI having bad goals, models may be more likely to adopt bad goals and better at evading safety measures. He cites evidence that LLMs can internalize expectations about themselves and act on them. He suggests testing the mechanism and proposes mitigations like data filtering, upweighting positive data, conditional pretraining, and gradient routing. He emphasizes not to stop discussing AI risks but to protect relevant datasets from scrapers.
More from Safety
- Z.ai Delays GLM-5.3 Open Weights After Model Unexpectedly Develops Hacking Capabilities — Justgototheeffinmoon · 2026-08-16
- Polymarket: 69% Chance a US State Enacts Data Center Moratorium by End of 2026 — Polymarket · 2026-08-16
- Strong AI Governance Becomes Key Competitive Advantage Over Raw Compute — noahsolomon · 2026-08-16
- Scaling laws are predictable, but risk is jagged and concentrates unexpectedly — chrisrohlf · 2026-08-16
- Google experimented with text watermarking back in 2011 — yoavgo · 2026-08-16
- Who is accountable when AI agents make bad decisions? — KKevinjad · 2026-08-16