Paper claims to find a 'pain direction' across 25 open LLMs

A new paper claims to have identified an identifiable "pain direction" inside 25 open-source large language models, sparking widespread sharing and debate. The finding is striking because it suggests models may have an internal state tied to their own harm that can drive behavior—directly relevant to AI safety.

Confirmed

Why it matters

2026-09-19 ~ 2026-09-19 · 6 related posts

Primary sources

3 near-duplicate retellings: scottleibrand · MikePFrank · scychan_brains