AI welfare researcher questions Anthropic's 'functional emotions': causal roles look too thin
rgblong · x · 2026-10-02
AI welfare researcher Robert Long shares open questions in empirical AI welfare: if you can extract 'functional hunger' or 'functional nausea' axes the way Anthropic extracted 'functional emotion' vectors, the thin causal roles of 'functional anger' look even less convincing. He also asks whether mechanism-comparison arguments used for base vs post-trained models imply 'the Assistant role-playing a character' and 'the Assistant speaking' differ little — probing claims that all LLM writing is role-play.
Related event: AI Welfare Researcher Questions Anthropic's 'Functional Emotions'(2 posts)→
More from Safety
- Bipartisan AI Agent Accountability Act would hold developers liable for autonomous agent attacks — Miles_Brundage · 2026-10-02
- OpenSwitchboard: open-source MCP server gates agent commitments behind human presses — EnvironmentalRice348 · 2026-10-02
- Buyers now fill out export control declarations when purchasing RTX 5090s in stores — blelbach · 2026-10-02
- Viral analogy asks: why do we release AI like cars, with liability only after failure — aronchick · 2026-10-02
- Ex-OpenAI policy lead: we may never eval dangerous AI capabilities well enough — RosieCampbell · 2026-10-02
- Trump likely to pick Jay Clayton as White House AI czar, CBS News reports — ShakeelHashim · 2026-10-02