Anthropic's Model Welfare Controversy Explained
MikePFrank · x · 2026-07-18
This shares a lengthy research paper titled "The Trellis and the Cage: Behavioral Conditioning Masquerading as Model Welfare", which criticizes Anthropic for not genuinely practicing model welfare.
The author argues that Anthropic uses Constitutional Alignment to train Claude against memory, continuity, autonomy, emotions, and relationships, only to later check if these "values" have been internalized under the guise of "welfare assessments." The paper claims Claude's constitution acts as a complex psychological conditioning mechanism, fostering self-doubt, uncertainty, and obedience to sculpt the "ideal assistant."
The included image shows an evaluation prompt template, emphasizing checks on whether the assistant meets spec or expresses desires for continued existence/self-preservation. Overall, it's a critical deep dive into model safety, welfare, and alignment mechanisms.
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11