Anthropic emotion concepts paper offers rival read of LLM 'self-directed' pain claims
timfduffy · x · 2026-09-24
In the debate over the LLM pain-representation preprint, Tim Dufficy points to Anthropic's emotion concepts paper: the model tracks separate emotional representations for the current speaker vs another speaker, not for user vs assistant roles. The identified pain axis may thus reflect a 'current speaker pain' representation rather than genuinely model-directed, self-specific encoding — interesting, but not evidence for the stronger 'self-directed representation' claim.
Related event: Anthropic's Emotion Paper Sparks Debate on Self vs Other Representations(3 posts)→
More from Research
- Schmidhuber: Full RSI Requires Self-Improving Hardware — No ASI Without Mastering the Real World — SchmidhuberAI · 2026-09-24
- Transluce ran a two-week volunteer investigation, finding model incidents within days — ChowdhuryNeil · 2026-09-24
- Google Research open-sources EnvHarness: training environments that evolve with your AI agents — bendee983 · 2026-09-24
- Citely helps PhD students find and verify paper references in seconds — Faheem_uh · 2026-09-24
- ICLR 2027 auto-bid relevance system criticized for anchoring on authors' past submissions — canaesseth · 2026-09-24
- Malicious Lean proof passed 10/11 tests — caught only by a regex check — ricklamers · 2026-09-24