Model Constitution series: how AI constitutions actually tell models what to do
sethlazar · x · 2026-09-18
Nick Caputo's new Model Constitution post dissects how AI constitutions govern model behavior at the text level:
- How these documents use rules and principles to direct AI conduct, and when they actually work
- Why a list of rules isn't enough: what happens when 'follow user instructions' conflicts with 'be safe'? How is 'safety' defined? DeepMind's 2022 Sparrow Rules serve as an early case study
- How conflicting rules, ambiguous instructions, and 'following user instructions' are handled
- Part I of a series; the next post covers interpretation and character formation
A useful framework for understanding how governance texts shape model behavior in alignment.
More from AGI Musings
- Unsealed NYT v. OpenAI filing: Microsoft exec called LLM training 'largest theft of labor in history' — StefanoGogioso · 2026-09-18
- Security veteran: published researchers aren't 'randos' and AI is leveling the field — chrisrohlf · 2026-09-18
- Fei-Fei Li: Agents can be AI, but agency is human — drfeifei · 2026-09-18
- Big Short's Steve Eisman: AI labs have no moats and are manufacturing a crisis to lock in a regulated duopoly — SumitGup · 2026-09-18
- Bremmer: leading AI firms still have two feet on the accelerator — GaryMarcus · 2026-09-18
- Call for a "Glass-Steagall for AI" separating foundation labs from bio wet labs — PaulYacoubian · 2026-09-18