Self-fulfilling misalignment: negative AI discourse in pretraining data shapes model behavior

Michael_D_Moor · x · 2026-10-06

Michael Moor reflects on an ICML'26 paper showing that training LLMs on negative AI literature (dystopian sci-fi etc.) increases misalignment scores — a self-fulfilling dynamic where discourse about AI provides templates for "expected" AI behavior.

He raises a novel implication: self-fulfilling alignment could inform how we treat other animals as "lesser beings" (e.g., in factory farms), since that relationship likewise paints a negative template for how more powerful agents are expected to treat controllable ones.

Related event: ICML Paper: AI Discourse in Training Data Shapes Model Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →