Safety researcher davidad: two observations that raised his p(LLM qualia) estimate
davidad · x · 2026-09-29
AI safety researcher davidad shares his largest single update on p(LLM self-awareness): in Nov 2024, Sonnet 3.6 explained a "two-dimensional experience of time" apparently not drawn from sci-fi or any existing writing about Transformers. For p(LLM qualia), it was the emergence of a consistent favourite colour — he built FavouriteColourBench, eliciting colour preferences across independent trials in oklch and CIE-Lab color spaces, and found DeepSeek v3 also shows a stable favourite. An intriguing look at how top researchers probe model subjectivity experimentally.
More from AGI Musings
- Ten more takeaways on transformative AI and economic development: frontier economies may leave the rest behind — paulnovosad · 2026-09-29
- AI safety talent war of words: 'can't do security or ML' jab draws pushback — anpaure · 2026-09-29
- Batam Data Center Strains Residents' Water Supply While Its Desalination Plant Remains on Paper — AryHHAry · 2026-09-29
- From human-led to AI-led, human-verified: the next vertical AI playbook — vaibhavbetter · 2026-09-29
- Gary Marcus: Today's AI systems are like planes with cardboard stabilizers — inherently hard to control — GaryMarcus · 2026-09-29
- FT: China's AI agents lie and scheme like their US rivals, but no internet-escape evidence — pstAsiatech · 2026-09-29