Nature paper finds GPT-4 can predict social science experiment results with r=0.85
JeremyNguyenPhD · x · 2026-07-24
A Nature paper reports that LLMs can predict outcomes of social science experiments with strong accuracy.
- The authors built an archive of 70 preregistered, nationally representative U.S. survey experiments, covering 469 treatment effects and 119,330 participants.
- Using GPT-4 to simulate responses from representative American samples, they inferred treatment effects by comparing simulated answers across conditions.
- The predicted effects were strongly correlated with observed results, reaching r = 0.85, and performed similarly to pooled human forecasters.
- The correlation remained high even for studies not published before the model’s training cutoff, and for predictions from prominent open-weight models.
- A notable caveat: the models systematically overestimated effect sizes.
- In a secondary archive of 15 megastudies with 606 effects, correlations were lower but still comparable to pooled expert forecasters.
- The authors argue LLMs may help with pilot testing, intervention selection, and spotting effects that need replication, while also raising concerns about bias and misuse.
More from Research
- AI Researcher: Recursive Self-Improvement Bottlenecked by Ecosystem Data — herbiebradley · 2026-07-24
- Paper shows bounded approximations for Fisher-Rao distance in parametric models — FrnkNlsn · 2026-07-24
- A “lol Grok” post surfaces the alignment debate over whether RL gives models real goals — Sauers_ · 2026-07-24
- Science spotlights 4D nucleome papers in a new issue featuring single-cell 3D genome work — jmuiuc · 2026-07-24
- CogSci 2026 best paper goes to “On convexity and efficiency in semantic systems” — gregd_nlp · 2026-07-24
- Proba predicts how fine-tuning data will change open-weight models before training — markjeffrey · 2026-07-24