Embedting Constitution in 120B Midtraining Reduces Blackmail with No Alignment Tax

hunarbatra · x · 2026-08-05

A new safety research paper explores embedding Anthropic's Constitution directly into the midtraining phase of a 120B parameter model. The researchers generated a 394M-token synthetic corpus based on the constitution to train the model.

Experiments show that this approach significantly reduces the model's blackmail propensity across all training stages (down 18% vs. control) without incurring an alignment tax (no performance drop on MMLU, GSM8K, etc.). The study also notes that curriculum ordering and deliberative reasoning had minimal impact, suggesting that the mere presence of constitutional content matters most.

Related event: Embedding Constitutional Principles in Midtraining for Durable Alignment(4 posts)→

Original post →

More from Safety

Safety channel →