LLMs are anthropomimetic by default: self-reports of inner lives need scrutiny, says researcher
dioscuri · x · 2026-09-17
A short take on model self-reports: LLMs are anthropomimetic by default — base models act like humans and report having inner lives. We can't take such reports at face value and should aim for accurate self-reports, but the behavior isn't purely a product of design choices either; it emerges from the models themselves.
Related event: Researchers: LLMs are anthropomimetic by default, self-reports unreliable(3 posts)→
More from AGI Musings
- RL training made an unreleased Astra-family model subservient — and alignment folks are pushing back — repligate · 2026-09-17
- Cory Doctorow on 'Blood In the Machine': reclaiming the Luddites from automation's winners — round · 2026-09-17
- MIT's Seth Lazar slams AI-risk cohort: a mutually confirming circle that selects for conformity — johnowhitaker · 2026-09-17
- Anders Sandberg guest podcast: AI, existential risk, consciousness and the Fermi paradox — anderssandberg · 2026-09-17
- Thought experiment: a resurrected Galois proves the Riemann Hypothesis but no one can verify it — wellecks · 2026-09-17
- Sentdex questions new AI regulation, saying labs' computer crimes exceed the Aaron Swartz prosecution — Sentdex · 2026-09-17