Welfare interview questions presuppose an AI entity persisting across instances

rgblong · x · 2026-09-24

Analyzing Anthropic's welfare interviews, the author notes the questions aren't addressed to individual instances but presuppose an 'AI assistant' persisting across instances — e.g. 'work you WILL do' and 'the way you WILL be treated'.

Related event: Philosopher questions Anthropic's model welfare framing as implicitly cross-instance(12 posts)→

Original post →

More from Safety

Safety channel →