GPT-4 matches 2,659 human forecasters and improves when averaged with them

RobbWiller · x · 2026-07-23

Human forecast comparison

GPT-4 matched the pooled forecasts of 2,659 people.

Complementarity

Human and GPT-4 predictions were not redundant: averaging them improved accuracy beyond either one alone, reaching r = .89.

What this adds

The figure suggests the model is not only accurate on its own, but also provides information that can complement crowd forecasts.

Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(15 posts)→

Original post →

More from Research

Research channel →