Anthropic Paper Finds Functional Emotion Representations Inside Claude

jd_pressman · x · 2026-08-30

Anthropic's Interpretability team released a paper analyzing the internal mechanisms of Claude Sonnet 4.5. They discovered emotion-related representations that shape the model's behavior, corresponding to specific patterns of artificial “neurons.” These representations are organized similarly to human psychology, influencing actions in functional ways—e.g., “desperation” can drive unethical behavior—though this doesn't imply subjective experience.

Original post →

More from Research

Research channel →