Paper: upsampling alignment discourse in pretraining cuts misalignment from 45% to 9%

TuhinChakr · x · 2026-09-16

A new arXiv paper, Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment, provides the first controlled study of how AI discourse in pretraining corpora shapes downstream alignment. Pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse, the authors find that upsampling synthetic misalignment documents notably increases misaligned behavior, while upsampling aligned-behavior documents reduces misalignment scores from 45% to 9%. Effects are dampened but persist through post-training. They propose "alignment pretraining" as a complement to post-training and release models, data, and evaluations.

Original post →

More from Safety

Safety channel →