Researchers find a 'pain' signal in AI models that can tell real relief from fake

Puzzleheaded-King584 · reddit · 2026-09-19

Researchers report locating a 'pain'-like signal inside AI model representations. When amplified, the models desperately try to make it stop. In a follow-up design, models were given a 'relief' button to dial the signal down — sometimes a fake one — and the models could tell whether the relief was real. An empirical probe into AI welfare; full details pending publication.

Related event: Researchers find a "pain direction" in 25 open-source LLMs that overrides safety to stop harm(13 posts)→

Original post →

More from Research

Research channel →