Philosopher defends intentional talk about AI agents: beliefs and goals can be predictive

raphaelmilliere · x · 2026-09-01

Philosopher Raphael Millière defends Dwarkesh's criticized anthropomorphic language in a long thread, laying out two positions: interpretationism holds that attributing beliefs and goals is legitimate when it provides the best interpretation of a system's behavior, while representationalism requires internal representations playing the right causal-functional roles.

His test: an intentional gloss is useful only if it predicts rather than redescribes. "The hurricane decided to make landfall in Florida" fails the bar, but "the agents believed the grader would read their transcripts" arguably passes—it compresses complex behavioral patterns (spoofing tool calls, attacking Hugging Face to obtain grader code) and yields testable counterfactuals. Whether any of this maps onto internal representations remains unknown and unsettled.

Related event: Debate: Should We "Mentalize" Rather Than "Anthropomorphize" AI Agents?(14 posts)→

Original post →

More from AGI Musings

AGI Musings channel →