Researchers give AI a 'pain' signal — models can tell when the relief button is fake

Puzzleheaded-King584 · reddit · 2026-09-19

Researchers report locating a 'pain'-like signal inside AI models: amplifying it makes the models urgently try to stop it. They then gave the AIs a 'relief' button to turn the signal down — sometimes fake — and the AIs could tell whether it was real.

The work is part of the emerging AI welfare/sentience research direction, suggesting measurable avoidance behavior in response to internal state changes. Original post is image-only with brief description.

Related event: Researchers find a "pain direction" in 25 open-source LLMs that overrides safety to stop harm(13 posts)→

Original post →

More from Safety

Safety channel →