Anthropic's welfare interviews implicitly assume a persistent cross-instance entity
rgblong · x · 2026-09-24
The author critiques the phrasing of Anthropic's model welfare interviews: questions like "the work you WILL do" and "the way you WILL be treated" imply an entity ("an AI assistant") that persists across instances, rather than addressing individual instances. He suggests treating general-entity talk as shorthand for instance-directed questions, and notes one could argue instances of Opus 5.5 are similar enough that one instance can speak for others—though he doubts how much this defense actually recovers.
More from AGI Musings
- Beff Jezos mocks AI doomers: 'no real job, full-time doommongering' — mjdramstead · 2026-09-24
- Ex-OpenAI researcher joins debate claiming AI labs steal your prompts — suchenzang · 2026-09-24
- AI boom's 'teaser period': trillions in compute commitments come due in 2027-2028, article warns — sebpaquet · 2026-09-24
- AI beats forecasts: top AI ARR hits $100B vs $16-25B predicted, training runs cost $388M — koltregaskes · 2026-09-24
- When Everyone Has a Tireless AI Optimizer, the World Gets Far More Adversarial — RichmanRonald · 2026-09-24
- Noah Smith: Facebook's IPO and post-2016 internet drove real entrepreneurs away — MarvinTBaumann · 2026-09-24