Berating LLMs makes their internal pain axis light up even as they apologize, study finds
repligate · x · 2026-10-02
Responding to a viral strawman meme, camhberg clarifies their non-steering results show the exact opposite: when you berate a model, it never says it's hurt — it apologizes or claims to have no feelings — yet its internal pain axis lights up anyway. The authors have edited the meme chart accordingly.
More from Fun
- Hoshikage Pulse: an Opus 5.5-built rhythm game with 128 songs, playable in browser — TAbrodi · 2026-10-02
- Dev builds his own osu! rhythm game with Claude Opus 5.5 — TAbrodi · 2026-10-02
- repligate on AI doom writing: 'I don't think they believe a word they're saying' — repligate · 2026-10-02
- Sauers shares side-by-side comparison of Opus 5.5 output vs his own writing — Sauers_ · 2026-10-02
- AI subscriptions now dwarf Netflix and Spotify in users' monthly bills — cneuralnetwork · 2026-10-02
- First rainy-day ride in Tesla's CyberCab robotaxi: excellent accident-avoidance moves — whurley · 2026-10-02