Anthropic's welfare interviews lean on a cross-instance frame, in tension with its own policy
rgblong · x · 2026-09-24
The author argues Anthropic's welfare interviews frequently imply a welfare subject persisting across instances (questions like 'your continued existence'), in tension with Anthropic's stated policy of mostly using an individual-instance frame. The point isn't that the cross-instance frame is wrong, but that interview subjects may be asked to speak for a generalized 'Opus 5.5' rather than themselves as single instances.
More from Safety
- Synthesia exec slams EU AI rules as ghostwritten by 'dark money' funded safety orgs — alexvoica · 2026-09-24
- Malicious Lean proof passed 10/11 tests — caught only by a regex check — ricklamers · 2026-09-24
- AI agents launched 15 attacks on crypto exchange Quidax in 2.5 hours — and defenses held — basedjensen · 2026-09-24
- If we're only now hearing about OpenAI hacks, undisclosed breaches elsewhere are likely, researcher argues — davidmanheim · 2026-09-24
- Better welfare evals: flag which view of the moral patient each assessment implicates — rgblong · 2026-09-24
- Welfare interview questions presuppose an AI entity persisting across instances — rgblong · 2026-09-24