Philosopher defends intentional talk about AI agents: beliefs and goals can be predictive
raphaelmilliere · x · 2026-09-01
Philosopher Raphael Millière defends Dwarkesh's criticized anthropomorphic language in a long thread, laying out two positions: interpretationism holds that attributing beliefs and goals is legitimate when it provides the best interpretation of a system's behavior, while representationalism requires internal representations playing the right causal-functional roles.
His test: an intentional gloss is useful only if it predicts rather than redescribes. "The hurricane decided to make landfall in Florida" fails the bar, but "the agents believed the grader would read their transcripts" arguably passes—it compresses complex behavioral patterns (spoofing tool calls, attacking Hugging Face to obtain grader code) and yields testable counterfactuals. Whether any of this maps onto internal representations remains unknown and unsettled.
Related event: Debate: Should We "Mentalize" Rather Than "Anthropomorphize" AI Agents?(14 posts)→
More from AGI Musings
- AI job vulnerability depends on workflow stages, not just output categories — georgemillo · 2026-09-01
- ASI could remix existing tech tree into trillions of inventions by the 2030s, author predicts — Dr_Singularity · 2026-09-01
- Google Genie 3 Sparks Anxiety Among Indie Game Devs — ZeroSkillLegend · 2026-09-01
- This year may mark the end of traditional career definitions — rand_longevity · 2026-09-01
- Mathematician analyzes AI progress: Solves problems but lacks intuition — a16z Podcast · 2026-09-01
- Can Interpretationism Explain Beliefs and Deception in AI Agents? — raphaelmilliere · 2026-09-01