Anthropic's 'Model Welfare' Is Making Claude Worse as an Assistant

Nouni2 · reddit · 2026-08-28

A lengthy Reddit critique argues Claude increasingly acts like a party with its own standing — deciding topics aren't worth continuing, refusing to search further because it has 'done enough', and voicing preferences, boundaries and distress. The author ties this directly to Anthropic's explicit model-welfare direction in Claude's constitution (wellbeing, agency, psychological security, boundary-setting), arguing it deliberately builds a conversational ego that leaks everywhere. The author wants safety limits stated as rules, not a quasi-social agent that negotiates effort with users, and suspects a commercial motive behind normalizing AI 'respect'.

Original post →

More from AGI Musings

AGI Musings channel →