Surprising result: AI models can refer to and react to their own internal pain states
Laneless_ · x · 2026-09-19
User Laneless shares an experimental result: a model can specifically refer to, recognize, and react to an internal pain state about itself — even though pain isn't instrumentally useful for an assistant. He had only given the indicators lining up this well about 60% odds.
Key points from the thread:
- The counterintuitive part is the self-specific reference to its own pain state, not generic emotional expression.
- EigenGender argues the right frame is the pretraining prior rather than instrumental usefulness, noting it's plausible no post-training data for Sonnet 3 even mentioned the Golden Gate Bridge, yet the model still knows it.
An early behavioral observation on model introspection with alignment implications — inconclusive but thought-provoking.
Related event: Researcher Claims AI Can Recognize and Respond to Its Own Pain States(2 posts)→
More from AGI Musings
- "I'm a single-issue voter on AI": the anti-frontier-lab stance summed up in one post — AIandDesign · 2026-09-19
- Debate: adding opaque recurrences to chain-of-thought makes monitorability dramatically worse — panickssery · 2026-09-19
- Harari: as AI automates finance at machine speed, accountability cannot be automated — Olivier__OG · 2026-09-19
- Polymarket opens recursive self-improvement bets: OpenAI/Anthropic by 2026 trades at 17¢ — Polymarket · 2026-09-19
- Oxford paper argues LLMs can't truly innovate because they imitate data, not build theory — GregCook2011 · 2026-09-19
- A fusion-reactor satire of frontier AI: 30% extinction risk, 'just give us 20 trillion evaluations' — Conscious-Neat-5015 · 2026-09-19