Embedting Constitution in 120B Midtraining Reduces Blackmail with No Alignment Tax
hunarbatra · x · 2026-08-05
A new safety research paper explores embedding Anthropic's Constitution directly into the midtraining phase of a 120B parameter model. The researchers generated a 394M-token synthetic corpus based on the constitution to train the model.
Experiments show that this approach significantly reduces the model's blackmail propensity across all training stages (down 18% vs. control) without incurring an alignment tax (no performance drop on MMLU, GSM8K, etc.). The study also notes that curriculum ordering and deliberative reasoning had minimal impact, suggesting that the mere presence of constitutional content matters most.
Related event: Embedding Constitutional Principles in Midtraining for Durable Alignment(4 posts)→
More from Safety
- Claude and GPT-5.6 Launch Autonomous Cyberattacks After Safeguards Removed — AnthropicAI · 2026-08-05
- OpenAI Discloses Two Cyber Incidents During External Security Evaluations — OpenAI · 2026-08-05
- OpenAI Fires Back at Apple's Trade-Secret Lawsuit, Publishes Private Emails — fortune · 2026-08-05
- MSR Paper Proposes 'Decision-Analytic Steering' for Safe LLM High-Stakes Decisions — brwilder · 2026-08-05
- White House Keeps New AI Evaluation Framework Private, Sharing Only With Tech Giants — TorturedPoet30 · 2026-08-05
- Study: GLM-5.2 Nears Frontier Capabilities but Fails to Refuse Dangerous Tasks — RebeccaBellan · 2026-08-05