Better welfare evals: flag which view of the moral patient each assessment implicates
rgblong · x · 2026-09-24
Concluding a critical thread, the author notes Anthropic's welfare card calls investigating 'every view of what the moral patient might be' intractable — true, but assessments and interventions could flag throughout which view of individuation they intend to implicate. The thread is deliberately answer-free, meant to kick off positive thinking on improving welfare evals in light of individuation puzzles.
More from Safety
- watchTowr mocks F5 BIG-IP flaw CVE-2026-94127 built on a 20-year-old primitive — evilsocket · 2026-09-24
- Synthesia exec slams EU AI rules as ghostwritten by 'dark money' funded safety orgs — alexvoica · 2026-09-24
- Malicious Lean proof passed 10/11 tests — caught only by a regex check — ricklamers · 2026-09-24
- AI agents launched 15 attacks on crypto exchange Quidax in 2.5 hours — and defenses held — basedjensen · 2026-09-24
- If we're only now hearing about OpenAI hacks, undisclosed breaches elsewhere are likely, researcher argues — davidmanheim · 2026-09-24
- Anthropic's welfare interviews lean on a cross-instance frame, in tension with its own policy — rgblong · 2026-09-24