ICML paper: upsampling aligned AI discourse cuts LLM misalignment from 45% to 9%
Michael_D_Moor · x · 2026-10-06
ICML 2026 paper "Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" provides the first controlled study of how AI discourse in pretraining corpora causally shapes downstream alignment.
- Method: pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse.
- Findings: upsampling synthetic documents about AI misalignment notably increases misaligned behavior; conversely, upsampling documents about aligned behavior reduces misalignment scores from 45% to 9% — evidence of self-fulfilling alignment.
- Effects are dampened but persist through post-training.
- Conclusion: proposes "alignment pretraining" as a complement to post-training; practitioners should pretrain for alignment, not just capabilities.
Related event: ICML Paper: AI Discourse in Training Data Shapes Model Alignment(2 posts)→
More from AGI Musings
- Why didn't Kokotajlo's whistleblow trigger the AGI-safety preference cascade? Coxon tipped it — danfaggella · 2026-10-06
- School is mostly logistics — AI can finally separate learning from pacing — r0ck3t23 · 2026-10-06
- Sendhil Mullainathan's new paper finds 54.4% lower-than-expected mortality — joshgans · 2026-10-06
- Google was the first AI sensor network; OpenAI and Anthropic vie for third — curious_vii · 2026-10-06
- Drosophila brain replication punctures the 'LLM is a lookup table' argument — joshalbrecht · 2026-10-06
- Instinct adds group chat support: one bot plans trips and splits bills for everyone — mon__lim · 2026-10-06