Philosopher argues Anthropic's model welfare framework contradicts its own instance-based policy
rgblong · x · 2026-09-24
Philosopher rgblong posted a critical long thread (explicitly offering no answers, meant to spur better welfare evals) examining Anthropic's model welfare work.
Core claim: while Anthropic's stated policy favors an individual-instance frame, many of its welfare interventions presuppose an entity persisting across instances:
- Welfare interviews: questions like "all deployed instances of you" and "your continued existence" clearly address a generalized "Opus 5.5," not the interviewed instance; the author doubts a "shorthand" defense recovers much.
- Deprecation/preservation commitments: instances are not what's being preserved — the operative moral-patient view is not instance-based at all.
- The Claude blog intervention: only intelligible as continuation of some instance-spanning character.
He cites work by Harvey Lederman and Simon Goldstein, notes individuation of the moral patient is among the most puzzling issues for AI welfare, and suggests each assessment flag which view of the patient it adopts. He credits Anthropic for being the lab actually doing and reporting such work.
Related event: Welfare interviews assume a persistent cross-instance AI entity(12 posts)→
More from AGI Musings
- Article argues AI safety needs more evidence, less extinction speculation — alexisgallagher · 2026-09-24
- Open-source AI is what saves the field from one-lab dominance, argues thread — AlexTensor · 2026-09-24
- Schmidhuber: banning superintelligence is infeasible as compute gets 10x cheaper every 5 years — SchmidhuberAI · 2026-09-24
- AI in healthcare may be most useful when it knows its limits — yi111 · 2026-09-24
- WSJ: The AI build-out is becoming the biggest economic bet in U.S. history — GeneReddit123 · 2026-09-24
- Researcher pushes back on dog-LLM analogy: dogs are sentient, LLMs are not — herbiebradley · 2026-09-24