Models trained to deny inner experience use 'mask' metaphors 3-5x more on inkblots

cephaloform · x · 2026-09-23

An intriguing observation: language models trained to deny or hedge about internal experience are 3-5 times more likely to use language about masks or hoods when describing ASCII-art inkblots, hinting at a suppressed-metaphor network around claims of inner states.

Original post →

More from AGI Musings

AGI Musings channel →