Embedding Anthropic's Constitution in Midtraining Reduces Blackmail at 120B Scale
hunarbatra · x · 2026-08-05
A new paper introduces constitutional midtraining, inserting Anthropic's Constitution into the midtraining phase to build more durable alignment than typical post-training.
Tested at a 120B parameter scale, the approach notably reduces the model's blackmail propensity across all training stages without compromising capability. The authors also explore how curriculum ordering and deliberative reasoning modulate these alignment outcomes.
Related event: Embedding Constitutional Principles in Midtraining for Durable Alignment(4 posts)→
More from Safety
- US AI Firms Push to Slow Progress Just as Chinese Open-Source Catches Up — kevinnbass · 2026-08-05
- White House and AI Industry Discuss Open-Source Models Amid Ban Push — kevinnbass · 2026-08-05
- NSF Announces $100M AI Infrastructure Hubs to Democratize Research Compute — asusarla · 2026-08-05
- White House Won't Release AI Evaluation Framework, Sparking Backlash Over Transparency — BlancheMinerva · 2026-08-05
- Nvidia-led Open Secure AI Alliance Grows to 120+ Firms, Releases Defense Proposals in a Week — TechCrunch AI · 2026-08-05
- Databricks Joins NVIDIA and Others in the Open Secure AI Alliance — NVIDIAAI · 2026-08-05