Researchers question whether AI pain probes measure model pain minus user pain
rgblong · x · 2026-09-23
In a technical debate about interpretability probes for model "pain", @agitbackprop suggests the probe may be a measurement artifact: depending on how it's fit, it could capture (model pain - user pain) rather than the model's own pain. rgblong adds that "negative world state", used as one of the control categories, oddly includes sentences about other people suffering, flagging potential confounds in the probe design.
More from Research
- Raw LLM probabilities aren't enough for decisions — calibration matters, researchers argue — PMinervini · 2026-09-23
- Parallel search blunts Grover's speedup, making AES-256 harder to break than thought — Jsevillamol · 2026-09-23
- Active learning splits composition from processing in optical materials optimization — bravo_abad · 2026-09-23
- Hodge and Yang-Mills still lack full Lean formalizations, making near-term solutions less likely — Jsevillamol · 2026-09-23
- Steve Hsu: AI will push math frontier far beyond human minds, compression defines 'human math' — burny_tech · 2026-09-23
- AI boosts science productivity — but mostly spawns spam papers, says Ehud Reiter — EhudReiter · 2026-09-23