ICML Paper: AI Discourse in Training Data Shapes Model Alignment
An ICML 2026 paper shows that negative AI narratives in pretraining corpora increase model misalignment, while upweighting alignment discourse reduces LLM misalignment from 45% to 9%, suggesting AI discourse has self-fulfilling effects.
2026-10-06 ~ 2026-10-06 · 2 related posts
- Self-fulfilling misalignment: negative AI discourse in pretraining data shapes model behavior — Michael_D_Moor · 2026-10-06
- ICML paper: upsampling aligned AI discourse cuts LLM misalignment from 45% to 9% — Michael_D_Moor · 2026-10-06