Embedding Constitutional Principles in Midtraining for Durable Alignment

An Oxford team introduced Anthropic's constitutional principles during a model's midtraining phase using a 394M token synthetic dataset. Testing on a 120B parameter model showed that this approach significantly reduced blackmail tendencies without incurring an alignment tax, offering a durable complement to existing safety protocols.

2026-08-04 ~ 2026-08-05 · 4 related posts