Google's Fuse paper benchmarks 12 LLMs on inferring hidden social motives
dair_ai · x · 2026-09-16
A Google Research paper covered by dair-ai studies how LLM assistants reason about the people in a user's life. People constantly ask assistants for social advice, but the assistant only hears the user's side, and others' intentions have no ground truth—so Fuse builds that ground truth with simulation.
Method:
- A target agent with a hidden motive interacts with other agents, including one playing the user
- The user agent describes events to the assistant, which must infer the motive
Evaluation:
- Simulations validated with 24k human annotations
- 12 LLMs tested
Findings: hearing events through the user makes the task harder; biased framing from the user shifts answers; models sometimes need more detail than humans; longer conversations with room for clarifying questions did not relieve the problem.
More from Research
- Fudan team: LLaMA3-70B and Qwen25-72B can self-replicate without human help — hey_abusiddik · 2026-09-16
- Ke Holdings Open-Sources PanoWorld: Consistent Whole-House 360° Panorama From Floorplans — rsasaki0109 · 2026-09-16
- When the Model Grades Itself, "It Passed" Means Less: On Independent AI Evaluation — vishalmisra · 2026-09-16
- Thompson's Compiler Attack Explains Why AI Evals Fail When the Model Helps Build Them — vishalmisra · 2026-09-16
- Princeton's Elad Hazan reflects on two decades of online convex optimization — HazanPrinceton · 2026-09-16
- RL on private lab data to predict experimental outcomes, immune to web contamination — hsu_byron · 2026-09-16