ICML Paper: AI Discourse in Training Data Shapes Model Alignment

An ICML 2026 paper shows that negative AI narratives in pretraining corpora increase model misalignment, while upweighting alignment discourse reduces LLM misalignment from 45% to 9%, suggesting AI discourse has self-fulfilling effects.

2026-10-06 ~ 2026-10-06 · 2 related posts