AI welfare researcher questions Anthropic's 'functional emotions': causal roles look too thin

rgblong · x · 2026-10-02

AI welfare researcher Robert Long shares open questions in empirical AI welfare: if you can extract 'functional hunger' or 'functional nausea' axes the way Anthropic extracted 'functional emotion' vectors, the thin causal roles of 'functional anger' look even less convincing. He also asks whether mechanism-comparison arguments used for base vs post-trained models imply 'the Assistant role-playing a character' and 'the Assistant speaking' differ little — probing claims that all LLM writing is role-play.

Related event: AI Welfare Researcher Questions Anthropic's 'Functional Emotions'(2 posts)→

Original post →

More from Safety

Safety channel →