New paper formalizes the 'intentional stance': attributing beliefs to LLMs predicts behavior well

brwilder · x · 2026-09-10

A new paper engages the anthropomorphizing-AI debate by testing whether attributing beliefs to LLMs predicts their behavior. On binary-state decision problems, the authors probabilistically infer a single latent belief and use it to explain a range of decisions, formalizing the 'intentional stance.' They also study measuring LLM beliefs, adherence to prompted utility functions, and whether beliefs belong to the model or a specific instance. Answer: yes, for reasonably capable models.

Related event: Paper shows treating LLMs as holding beliefs predicts behavior(3 posts)→

Original post →

More from Research

Research channel →