Paper: upsampling alignment discourse in pretraining cuts misalignment from 45% to 9%
TuhinChakr · x · 2026-09-16
A new arXiv paper, Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment, provides the first controlled study of how AI discourse in pretraining corpora shapes downstream alignment. Pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse, the authors find that upsampling synthetic misalignment documents notably increases misaligned behavior, while upsampling aligned-behavior documents reduces misalignment scores from 45% to 9%. Effects are dampened but persist through post-training. They propose "alignment pretraining" as a complement to post-training and release models, data, and evaluations.
More from Safety
- Bitsec's multi-model agent stack found 160+ exploits, beating a single 'superhuman' model — markjeffrey · 2026-09-16
- Podcast: Oxford's Carissa Véliz on Meta's landmark lawsuit, surveillance and AI prediction — CarissaVeliz · 2026-09-16
- Pedro Domingos mocks EU AI Act as the only reason AI hasn't wiped out humanity — pmddomingos · 2026-09-16
- Investigation Claims EA Donors Funded Guardian's AI Coverage: All 6 Participants Paid by Same Ecosystem — beffjezos · 2026-09-16
- AI 2027 authors pitch Plan A: delay superintelligence to 2040 with fully open AI research — Turn_Trout · 2026-09-16
- EA's media capture and doomer headlines skew public AI perception, argues Nahom Sisay — NathanpmYoung · 2026-09-16