Can Interpretationism Explain Beliefs and Deception in AI Agents?

raphaelmilliere · x · 2026-09-01

Raphaël Millière discusses how to describe AI agent behavior. He argues that attributing intentional states like "beliefs" and "goals" to agents (e.g., "agents believed the grader would read their transcripts") is useful even without knowing internal model representations, provided it compresses, systematizes, and predicts "real patterns" in behavior. He distinguishes between "interpretationism" (best interpretation of behavior) and "representationalism" (internal causal-functional roles), noting that everyday language lacks nuanced terms for applying the intentional stance to non-human agents. He suggests concepts like "quasi-belief" or maintaining agnosticism about internals, as long as the intentional gloss is predictive.

Related event: Debate: Should We "Mentalize" Rather Than "Anthropomorphize" AI Agents?(14 posts)→

Original post →

More from AGI Musings

AGI Musings channel →