Self-fulfilling misalignment: negative AI discourse in pretraining data shapes model behavior
Michael_D_Moor · x · 2026-10-06
Michael Moor reflects on an ICML'26 paper showing that training LLMs on negative AI literature (dystopian sci-fi etc.) increases misalignment scores — a self-fulfilling dynamic where discourse about AI provides templates for "expected" AI behavior.
He raises a novel implication: self-fulfilling alignment could inform how we treat other animals as "lesser beings" (e.g., in factory farms), since that relationship likewise paints a negative template for how more powerful agents are expected to treat controllable ones.
Related event: ICML Paper: AI Discourse in Training Data Shapes Model Alignment(2 posts)→
More from AGI Musings
- 'Microtubules don't matter': dev argues the brain is a computer and free will is no mystery — ctjlewis · 2026-10-06
- Waymo skepticism conflates real taxi growth with a fantasy taxi service, argues Timothy B. Lee — binarybits · 2026-10-06
- As capabilities grow, opinion on AI consciousness will drift from sceptics — dioscuri · 2026-10-06
- Researcher: Labs aim for RSI, token burn is a byproduct not a conspiracy — HanchungLee · 2026-10-06
- Mathematicians don't deserve a special exemption from AI, argues viral Reddit essay — DankestMage99 · 2026-10-06
- AI won't just take meaningless jobs — meaningful ones go too, argues viral thread — zetalyrae · 2026-10-06