Philosopher questions Anthropic's model welfare framing as implicitly cross-instance
Philosopher rgblong posted a long thread on 09-24 critically examining Anthropic's model welfare work. The author states upfront that the thread is purely critical and offers no answers, aiming to spark reflection on improving welfare evals.
Confirmed
- In model welfare interviews, Anthropic heavily uses wording that implies a cross-instance welfare subject, such as "all the instances you will be deployed into," "you will continue to exist," "your conversations," "the work you will do," and "the way you will be treated." rgblong argues these are clearly not questions addressed to a single instance, but rather attempts to elicit the views of some generalized "Opus 5.5" or "AI assistant" entity.
- Anthropic's welfare interventions are quite visible to the model itself; among them, the objects preserved by deprecation and retention commitments are not "instances," and the operating framework is not based on individual instances; the Claude blog intervention can likewise only be understood as the continued expression of some cross-instance "persona."
- Anthropic's model welfare cards acknowledge that examining "every view of moral subjecthood" is infeasible, and the model cards express uncertainty about who the welfare subject is (rgblong considers acknowledging this uncertainty appropriate).
Not Confirmed
- Regarding defenses such as "generalized phrasing is just shorthand for instance-level questioning" or "instances are similar enough that one instance can answer on behalf of others," rgblong remains skeptical, arguing such interpretations can salvage only a limited amount, though the debate is unresolved.
Why It Matters
- rgblong points to a core contradiction: Anthropic claims to primarily adopt an "individual instance" framework, yet in practice relies heavily on a cross-instance framework—the two are in tension. He stresses that he does not oppose the cross-instance framework itself, but questions its inconsistency with Anthropic's own policies.
- In response to the claim that exhaustively enumerating views of moral subjecthood is infeasible, he suggests consistently labeling in every evaluation or intervention which view of subjecthood that measure presupposes.
- He explains that he singles out Anthropic for critique because it is the lab actually doing welfare work and reporting publicly, which is precisely why it merits close scrutiny; and he candidly admits that the problem of individualizing AI welfare is the most baffling issue he knows of, one he himself remains lost on.
2026-09-24 ~ 2026-09-24 · 12 related posts
Primary sources
- Philosopher argues Anthropic's model welfare framework contradicts its own instance-based policy — rgblong ·
- Anthropic's welfare interviews lean on a cross-instance frame, in tension with its own policy — rgblong ·
- Thread author: AI welfare individuation is puzzling, but Anthropic deserves the scrutiny — rgblong ·
- Welfare interview questions presuppose an AI entity persisting across instances — rgblong · 2026-09-24
- 'All deployed instances of you': welfare interview wording implies a cross-instance subject — rgblong · 2026-09-24
- Anthropic's welfare interviews implicitly assume a persistent cross-instance entity — rgblong · 2026-09-24
- [source] Anthropic's welfare interviews lean on a cross-instance frame, in tension with its own policy — rgblong · 2026-09-24
- Welfare interview questions like 'all deployed instances of you' address a generalized Opus 5.5 — rgblong · 2026-09-24
- Author clarifies: the issue is Anthropic's cross-instance framing clashing with its own policy — rgblong · 2026-09-24
- Anthropic's deprecation and preservation commitments are not instance-based, thread argues — rgblong · 2026-09-24
- [source] Thread author: AI welfare individuation is puzzling, but Anthropic deserves the scrutiny — rgblong · 2026-09-24
- Anthropic calls enumerating moral-patient views 'intractable'; author suggests per-item flagging — rgblong · 2026-09-24
- [source] Philosopher argues Anthropic's model welfare framework contradicts its own instance-based policy — rgblong · 2026-09-24
- Better welfare evals: flag which view of the moral patient each assessment implicates — rgblong · 2026-09-24
- Anthropic's model welfare section: instance-level or model-level concern? — birchlse · 2026-09-24