A 70-experiment benchmark finds GPT-4’s predicted effects at r=.85 versus reality

RobbWiller · x · 2026-07-23

Method and result

The authors evaluated GPT-4 on 70 preregistered experiments from TESS and replications, covering 469 effects and 119,330 participants.

They prompted the model with the study materials and demographic profiles, analyzed the simulated responses, and compared the inferred treatment effects with the real ones. The resulting estimates were highly correlated with the actual effects: r = .85, adjusted r = .92.

Broader pattern

Performance stayed strong across many social-science subfields, not just one narrow task.

Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(15 posts)→

Original post →

More from Research

Research channel →