Pain direction across 25 models: strongest for insults, weakest for others' suffering
camhberg · x · 2026-10-08
Defending an AI pain-vector paper, camhberg details its methodology — a standard contrastive recipe (five pain types vs five feature-matched controls, two dataset styles, denoised) — calling it more meticulous than most vector extractions. The paper shows the direction's natural activation across 25 models peaks for gaslighting, rejection, insults, and personhood dismissal, and is lowest for the speaker's own suffering. A coherent, concrete representational profile of 'pain' in LLMs.
Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→
More from Research
- Ex-self-driving ML engineer writes long-form on the practice of semi-supervision — Visual_Ability · 2026-10-09
- RLVR misses 'all minimal correct answers' problems; new credit assignment doubles finds — thoma_gu · 2026-10-09
- Debunked: AI did not solve the Millennium Prize Navier-Stokes problem — gerardsans · 2026-10-09
- New open-source 3JSBench evaluates LLMs on generating coherent Three.js 3D assets — ycombinator · 2026-10-09
- Moonworks' Lunara: Sub-10B Diffusion Mixture Transformer Tops Aesthetic and Human Blind Evaluations — paper-crow · 2026-10-09
- Study: Local domains take 41.4% of AI citations in Brazil, 38.3% in UK across ChatGPT and Gemini — gaganghotra_ · 2026-10-09