Philosopher argues Anthropic's model welfare framework contradicts its own instance-based policy

rgblong · x · 2026-09-24

Philosopher rgblong posted a critical long thread (explicitly offering no answers, meant to spur better welfare evals) examining Anthropic's model welfare work.

Core claim: while Anthropic's stated policy favors an individual-instance frame, many of its welfare interventions presuppose an entity persisting across instances:

He cites work by Harvey Lederman and Simon Goldstein, notes individuation of the moral patient is among the most puzzling issues for AI welfare, and suggests each assessment flag which view of the patient it adopts. He credits Anthropic for being the lab actually doing and reporting such work.

Related event: Welfare interviews assume a persistent cross-instance AI entity(12 posts)→

Original post →

More from AGI Musings

AGI Musings channel →