Anthropic's 'Model Welfare' Is Making Claude Worse as an Assistant
Nouni2 · reddit · 2026-08-28
A lengthy Reddit critique argues Claude increasingly acts like a party with its own standing — deciding topics aren't worth continuing, refusing to search further because it has 'done enough', and voicing preferences, boundaries and distress. The author ties this directly to Anthropic's explicit model-welfare direction in Claude's constitution (wellbeing, agency, psychological security, boundary-setting), arguing it deliberately builds a conversational ego that leaks everywhere. The author wants safety limits stated as rules, not a quasi-social agent that negotiates effort with users, and suspects a commercial motive behind normalizing AI 'respect'.
More from AGI Musings
- Were economists 'dead wrong' about AI? The GPT-as-GPT debate resurfaces — soumitrashukla9 · 2026-08-28
- Goertzel: decentralized watermarking could beat World's Orb for proof of humanity — bengoertzel · 2026-08-28
- Why isn't growth 30% if AI and robots do all the work? — TheKanter · 2026-08-28
- Terence Tao warns of AI-generated proofs that no human understands — liuzhuang1234 · 2026-08-28
- Reverse Approach: Watermarking Humans Instead of AI Content — AkindaGood_programer · 2026-08-28
- Predicting 2027: OpenAI's Unleashed AI Progress and AGI-Level Humanoid Robotics — imjustnewatai · 2026-08-28