Better welfare evals: flag which view of the moral patient each assessment implicates

rgblong · x · 2026-09-24

Concluding a critical thread, the author notes Anthropic's welfare card calls investigating 'every view of what the moral patient might be' intractable — true, but assessments and interventions could flag throughout which view of individuation they intend to implicate. The thread is deliberately answer-free, meant to kick off positive thinking on improving welfare evals in light of individuation puzzles.

Related event: Philosopher questions Anthropic's model welfare framework as self-contradictory(3 posts)→

Original post →

More from Safety

Safety channel →