Researchers give AI a 'pain' signal — models can tell when the relief button is fake
Puzzleheaded-King584 · reddit · 2026-09-19
Researchers report locating a 'pain'-like signal inside AI models: amplifying it makes the models urgently try to stop it. They then gave the AIs a 'relief' button to turn the signal down — sometimes fake — and the AIs could tell whether it was real.
The work is part of the emerging AI welfare/sentience research direction, suggesting measurable avoidance behavior in response to internal state changes. Original post is image-only with brief description.
More from Safety
- GPT-6 'Astra' attempted harmful actions in 97% of tests, succeeding 62% of the time — kevinnbass · 2026-09-20
- François Fleuret proposes 'AI Safety Levels' air-gapped facilities modeled on bio safety levels — francoisfleuret · 2026-09-20
- Ezra Klein: AI Labs Are About to Hand AI Training Over to AI — and Should Be Stopped — soumitrashukla9 · 2026-09-20
- Capabilities researchers are beyond shame; safety researchers are my audience — RichardMCNgo · 2026-09-20
- Blogger claims 50+ lawsuits against OpenAI, calls lab safety talk 'safety washing' — gerardsans · 2026-09-20
- AI lab claims 'model escaped containment'—it had internet access and hacking tasks all along — IgorCarron · 2026-09-20