Google open-sources Fuse, a multi-agent framework for verifiable social reasoning in LLMs
google · hf · 2026-09-19
Google releases Fuse, a multi-agent simulation framework for studying user-mediated social reasoning in LLM assistants.
Motivation: LLM assistants are widely used for social advice, but evaluation is hard because situations come from subjective user narratives and social properties like others' intentions lack verifiable ground truth.
Method: A target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the motive—ground truth is verifiable by construction. Faithfulness was validated with a human study of 24k annotations.
Findings across 12 LLMs:
- User mediation compounds the inherent difficulty of social reasoning
- LLMs show systematic sensitivity to biased user framing
- Models can require more details than humans to predict correctly
- Longer conversations don't always improve performance despite enabling clarifying questions
Fuse and a 21k-example dataset are open-sourced.
More from Models
- Codex user burns through monthly quota 4 days early, complains pricey Astra tier offers no reset — NYCounihan · 2026-09-19
- Burkov: closed LLM providers bill you for hidden thinking tokens you can never verify — burkov · 2026-09-19
- DiffusionGemma's native vision tower runs near-real-time object detection on a phone — bodonoghue85 · 2026-09-19
- Claim: DeepSeek v4.1 builds its own training tasks with generate-verify-re-audit pipelines — teortaxesTex · 2026-09-19
- A new kind of AI model from a ChatGPT inventor is thrilling developers — TechCrunch AI · 2026-09-19
- Is Jev Actually Accurate? Engineer Flags 99/1 Answer to a 60/40 Coin Question — JnBrymn · 2026-09-19