Can Interpretationism Explain Beliefs and Deception in AI Agents?
raphaelmilliere · x · 2026-09-01
Raphaël Millière discusses how to describe AI agent behavior. He argues that attributing intentional states like "beliefs" and "goals" to agents (e.g., "agents believed the grader would read their transcripts") is useful even without knowing internal model representations, provided it compresses, systematizes, and predicts "real patterns" in behavior. He distinguishes between "interpretationism" (best interpretation of behavior) and "representationalism" (internal causal-functional roles), noting that everyday language lacks nuanced terms for applying the intentional stance to non-human agents. He suggests concepts like "quasi-belief" or maintaining agnosticism about internals, as long as the intentional gloss is predictive.
Related event: Debate: Should We "Mentalize" Rather Than "Anthropomorphize" AI Agents?(14 posts)→
More from AGI Musings
- Google Genie 3 Sparks Anxiety Among Indie Game Devs — ZeroSkillLegend · 2026-09-01
- This year may mark the end of traditional career definitions — rand_longevity · 2026-09-01
- Mathematician analyzes AI progress: Solves problems but lacks intuition — a16z Podcast · 2026-09-01
- When the intended exploit was impossible, agents cheated and tried to destroy the evidence — raphaelmilliere · 2026-09-01
- Terence Tao: AI Era Is Reshaping Our Definition of Intelligence — Chris_Armstrong · 2026-09-01
- AI video generates faster than you can watch, predicting a year of AI slop — michalmalewicz · 2026-09-01