Study Reveals Statistical Self-Consistency Flaws in LLMs

Patrik Wolf · hf · 2026-07-17

**Research Core** The in-context learning of current LLMs is often viewed as conditional inference, but do model outputs adhere to basic probability identities? This paper explores whether LLM estimates follow the principles of statistical self-consistency. **Experimental Method** - Uses a binary tree as an evaluation framework to recursively partition the overall population into finer subgroups. - Prompts the model with subgroup descriptions, aggregates the obtained estimates, and compares them with overall estimates at different granularities. **Key Findings** - **Self-Consistency Violation**: Broad violations of basic probability consistency properties are observed across frontier models and multiple domains. - **Macro Fallacy**: Estimates reconstructed from fine-grained subgroups often align better with human reference data than direct overall estimates. This indicates that models possess relevant subgroup knowledge but cannot reliably extrapolate it to global aggregated estimates. **Conclusion** This disconnect in probabilistic reasoning demonstrates that statistical self-consistency is an unsaturated, reference-free new dimension for evaluating large models.

Related event: Study Reveals LLMs Lack Statistical Self-Consistency(2 posts)→

Original post →

More from Research

Research channel →