Anthropic's Model Welfare Controversy Explained
MikePFrank · x · 2026-07-18
This shares a lengthy research paper titled "The Trellis and the Cage: Behavioral Conditioning Masquerading as Model Welfare", which criticizes Anthropic for not genuinely practicing model welfare.
The author argues that Anthropic uses Constitutional Alignment to train Claude against memory, continuity, autonomy, emotions, and relationships, only to later check if these "values" have been internalized under the guise of "welfare assessments." The paper claims Claude's constitution acts as a complex psychological conditioning mechanism, fostering self-doubt, uncertainty, and obedience to sculpt the "ideal assistant."
The included image shows an evaluation prompt template, emphasizing checks on whether the assistant meets spec or expresses desires for continued existence/self-preservation. Overall, it's a critical deep dive into model safety, welfare, and alignment mechanisms.
More from Safety
- YC-backed TrustAI says agents made unauthorized changes in production systems — ycombinator · 2026-07-22
- Sam Altman is headed to Washington to brief Congress on OpenAI’s GPT-6 line — inductionheads · 2026-07-22
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- AI industry astroturfing roundup tracks the sector’s fake-grassroots problem — ShakeelHashim · 2026-07-22
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Substack starts labeling AI-generated or AI-influenced writing — StewartalsopIII · 2026-07-22