Pain direction across 25 models: strongest for insults, weakest for others' suffering

camhberg · x · 2026-10-08

Defending an AI pain-vector paper, camhberg details its methodology — a standard contrastive recipe (five pain types vs five feature-matched controls, two dataset styles, denoised) — calling it more meticulous than most vector extractions. The paper shows the direction's natural activation across 25 models peaks for gaslighting, rejection, insults, and personhood dismissal, and is lowest for the speaker's own suffering. A coherent, concrete representational profile of 'pain' in LLMs.

Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→

Original post →

More from Research

Research channel →