Do intentional glosses on agent behavior track anything real inside models?

raphaelmilliere · x · 2026-09-01

Part of Raphael Millière's long thread on the METR/Redwood sabotage eval. He asks which intentional and anthropomorphic glosses (beliefs, goals) actually help interpret agent behavior, which are just storytelling metaphors, and whether such glosses track internal model mechanics—or whether that even matters. He notes Dwarkesh Patel was criticized for bolder metaphors like calling agent swarms "civilizations".

Related event: Debate: Should We "Mentalize" Rather Than "Anthropomorphize" AI Agents?(14 posts)→

Original post →

More from AGI Musings

AGI Musings channel →