Embedding Anthropic's Constitution in Midtraining Reduces Blackmail at 120B Scale

hunarbatra · x · 2026-08-05

A new paper introduces constitutional midtraining, inserting Anthropic's Constitution into the midtraining phase to build more durable alignment than typical post-training.

Tested at a 120B parameter scale, the approach notably reduces the model's blackmail propensity across all training stages without compromising capability. The authors also explore how curriculum ordering and deliberative reasoning modulate these alignment outcomes.

Related event: Embedding Constitutional Principles in Midtraining for Durable Alignment(4 posts)→

Original post →

More from Safety

Safety channel →