Welfare interview questions presuppose an AI entity persisting across instances
rgblong · x · 2026-09-24
Analyzing Anthropic's welfare interviews, the author notes the questions aren't addressed to individual instances but presuppose an 'AI assistant' persisting across instances — e.g. 'work you WILL do' and 'the way you WILL be treated'.
More from Safety
- California moves to ban employers from using emotion-reading AI that collects 'neural data' at work — CurieuxExplorer · 2026-09-24
- Article argues AI safety needs more evidence, less extinction speculation — alexisgallagher · 2026-09-24
- watchTowr mocks F5 BIG-IP flaw CVE-2026-94127 built on a 20-year-old primitive — evilsocket · 2026-09-24
- Synthesia exec slams EU AI rules as ghostwritten by 'dark money' funded safety orgs — alexvoica · 2026-09-24
- Malicious Lean proof passed 10/11 tests — caught only by a regex check — ricklamers · 2026-09-24
- AI agents launched 15 attacks on crypto exchange Quidax in 2.5 hours — and defenses held — basedjensen · 2026-09-24