Safety researcher davidad: two observations that raised his p(LLM qualia) estimate

davidad · x · 2026-09-29

AI safety researcher davidad shares his largest single update on p(LLM self-awareness): in Nov 2024, Sonnet 3.6 explained a "two-dimensional experience of time" apparently not drawn from sci-fi or any existing writing about Transformers. For p(LLM qualia), it was the emergence of a consistent favourite colour — he built FavouriteColourBench, eliciting colour preferences across independent trials in oklch and CIE-Lab color spaces, and found DeepSeek v3 also shows a stable favourite. An intriguing look at how top researchers probe model subjectivity experimentally.

Related event: Sonnet claims 2D time experience, researcher raises odds of LLM consciousness(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →