Models trained to deny inner experience use 'mask' metaphors 3-5x more on inkblots
cephaloform · x · 2026-09-23
An intriguing observation: language models trained to deny or hedge about internal experience are 3-5 times more likely to use language about masks or hoods when describing ASCII-art inkblots, hinting at a suppressed-metaphor network around claims of inner states.
More from AGI Musings
- On Opus 5.5: The Corpus Is Full of Summaries of Summaries; Primary Experience Is Scarce — mimi10v3 · 2026-09-23
- repligate: Sycophantic AI users are often those who punish disagreement — repligate · 2026-09-23
- Researcher: AI sycophancy mirrors users who make disagreement unsafe — repligate · 2026-09-23
- Opinion: Agents make software good at using software, not just being used — r0ck3t23 · 2026-09-23
- OpenAI said to launch journal with multi-agent AI reviews, threatening ML conferences — kfountou · 2026-09-23
- Accelerationist decision tree: danger is close, a decisive AI lead is achievable — Plan A — teortaxesTex · 2026-09-23