Constitutional Midtraining Yields Durable Alignment Gains at No Capability Cost
hunarbatra · x · 2026-08-05
The authors conclude that adding a modest amount of constitutional content during midtraining provides broad, durable alignment gains without sacrificing model performance on benchmarks (no alignment tax). This suggests constitutional midtraining is a complementary addition to standard safety post-training pipelines.
Fully Open-Sourced: The team has released all code, datasets, benchmarks, and 15 matched 120B parameter checkpoints.
Related event: Embedding Constitutional Principles in Midtraining for Durable Alignment(4 posts)→
More from Safety
- US AI Firms Push to Slow Progress Just as Chinese Open-Source Catches Up — kevinnbass · 2026-08-05
- White House and AI Industry Discuss Open-Source Models Amid Ban Push — kevinnbass · 2026-08-05
- NSF Announces $100M AI Infrastructure Hubs to Democratize Research Compute — asusarla · 2026-08-05
- White House Won't Release AI Evaluation Framework, Sparking Backlash Over Transparency — BlancheMinerva · 2026-08-05
- Nvidia-led Open Secure AI Alliance Grows to 120+ Firms, Releases Defense Proposals in a Week — TechCrunch AI · 2026-08-05
- Databricks Joins NVIDIA and Others in the Open Secure AI Alliance — NVIDIAAI · 2026-08-05