Frontier LLMs Fail at Counterfactual Reasoning: Best Score Just 64.6%
AryHHAry · x · 2026-09-01
A new paper from Zhejiang University, Alibaba, and City University of Hong Kong reveals that while LLMs can give convincing answers to "What If" counterfactual questions, the underlying causal chains are often incoherent.
Research Core:
- WhatIfBench: A benchmark of 220 open-ended, long-horizon, cross-domain counterfactual questions (STEM 113, HSS 79, Hybrid 28).
- PRISM Framework: Evaluates the reasoning process rather than just the final outcome. It converts model answers into causal graphs (event, state, mechanism) and assesses them via Process Metrics (causal chain validity) and Rubric Metrics (explanation accuracy).
Results:
Six frontier models were tested, with none achieving saturation:
- GPT-5.5: 64.62%
- Claude Opus 4.7: 59.11%
- Gemini 3.1 Pro Preview: 56.33%
- DeepSeek-V4-Pro: 55.12%
- Qwen3-Max: 52.26%
- GLM-5.1: 51.99%
Key Failure Modes:
- Causal gaps: Logical leaps within the reasoning chain.
- Long-horizon breakdown: Coherence degrades as reasoning steps increase.
- State inconsistency: Contradictory descriptions of entity states throughout the text.
More from Research
- ProtRL Upgrade: Adds Distributed Training and Custom Loss Support — ferruz_noelia · 2026-09-01
- ExploitGym system prompts reveal little of interest — voooooogel · 2026-09-01
- ECCV 2026: EquiFusion Enables Kinematics-Agnostic Human Motion Prediction — CSProfKGD · 2026-09-01
- LightFuse SOTA Relightable 3D Gaussian Reconstruction Beats Baseline by 9.74dB — janusch_patas · 2026-09-01
- Doubt cast on SWA effective receptive field math — giffmana · 2026-09-01
- Call for Papers: Synthetic Populations Research at Frontiers in AI — frederickaplan · 2026-09-01