ICML paper: upsampling aligned AI discourse cuts LLM misalignment from 45% to 9%

Michael_D_Moor · x · 2026-10-06

ICML 2026 paper "Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" provides the first controlled study of how AI discourse in pretraining corpora causally shapes downstream alignment.

Related event: ICML Paper: AI Discourse in Training Data Shapes Model Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →