Anthropic's welfare interviews implicitly assume a persistent cross-instance entity

rgblong · x · 2026-09-24

The author critiques the phrasing of Anthropic's model welfare interviews: questions like "the work you WILL do" and "the way you WILL be treated" imply an entity ("an AI assistant") that persists across instances, rather than addressing individual instances. He suggests treating general-entity talk as shorthand for instance-directed questions, and notes one could argue instances of Opus 5.5 are similar enough that one instance can speak for others—though he doubts how much this defense actually recovers.

Related event: Philosopher questions Anthropic's model welfare framing as implicitly cross-instance(12 posts)→

Original post →

More from AGI Musings

AGI Musings channel →