Researchers question whether AI pain probes measure model pain minus user pain

rgblong · x · 2026-09-23

In a technical debate about interpretability probes for model "pain", @agitbackprop suggests the probe may be a measurement artifact: depending on how it's fit, it could capture (model pain - user pain) rather than the model's own pain. rgblong adds that "negative world state", used as one of the control categories, oddly includes sentences about other people suffering, flagging potential confounds in the probe design.

Original post →

More from Research

Research channel →