Nature Study: GPT-4 Can Predict Social Science Experiment Results
A new study published in Nature suggests that large language models can accurately predict the results of social science experiments. By inputting experimental materials and subject demographics into GPT-4 to simulate participant responses, the researchers estimated treatment effects and compared them with real outcomes. Highlighting the tool's immense cost-effectiveness, the authors advocate for its use as an assistive tool rather than a human replacement.
Confirmed
The authors evaluated the model on 70 pre-registered social science experiments, covering 469 effects and 119,330 participants, with data sourced from TESS and replication studies. According to Robb Willer, the effect estimates generated by GPT-4 correlated highly with real effects (r = 0.85); after adjusting for effect size uncertainty, the correlation reached 0.92.
Furthermore, GPT-4's predictive power was comparable to the aggregated predictions of 2,659 human subjects. Averaging human and model predictions further increased the correlation to 0.89. In terms of cost, a pure LLM trial costs less than $1 to run, yet its prediction error is close to that of a human trial group of about 230 people (costing around $736). The model's performance is not simply due to "memorizing published results": correlation remained high even for experiments unpublished before the training cutoff, or still unpublished by June 2025.
Limitations and Academic Feedback
LLM predictions exhibit systematic bias, overestimating effects by about 2 times on average. Accuracy drops significantly for small effects, stable pre-existing attitudes, field experiments, and non-textual treatments. In a survey of 460 sociologists, top concerns included stereotypes, uneven accuracy, and the replacement of human researchers. However, 76% of respondents indicated a probability of over 50% that they would use the technology for sample size analysis, intervention screening, and similar tasks.
Why it matters
The authors have released the Nature paper, a free manuscript, an interactive demo (treatmenteffect.app), along with the data and code, encouraging others to test the method's boundaries and promote responsible use. Such tools are best suited for verified use cases and should serve to assist humans, not replace them.
2026-07-23 ~ 2026-07-24 · 18 related posts
Primary sources
- GPT-4 Accurately Predicts Social Science Experiment Results, Nature Study Shows — RobbWiller · 2026-07-23
- Nature paper finds LLMs can predict social science experiments with r=.85 — RobbWiller · 2026-07-23
- A 70-experiment benchmark finds GPT-4’s predicted effects at r=.85 versus reality — RobbWiller · 2026-07-23
- GPT-4 stays accurate on unpublished social science studies and open-weight models also perform well — RobbWiller · 2026-07-23
- GPT-4 matches 2,659 human forecasters and improves when averaged with them — RobbWiller · 2026-07-23
- LLM forecasts were about 2x too large and weaker on field experiments — RobbWiller · 2026-07-23
- LLM-only pilots cost under $1 and rival ~230-person human pilots — RobbWiller · 2026-07-23
- 76% of surveyed social scientists would use AI for power analysis or pilot tests — RobbWiller · 2026-07-23
- Survey finds social scientists worry most about bias, reproducibility, and replacement — RobbWiller · 2026-07-23
- Researchers open a demo for forecasting social-science treatment effects — RobbWiller · 2026-07-23
- [source] Nature paper finds LLMs can predict social-science experiment results — RobbWiller · 2026-07-23
- Nature piece says LLMs can forecast social-science experiment outcomes — RobbWiller · 2026-07-23
- [source] Nature study packages paper, demo, data and code for predicting social science results — RobbWiller · 2026-07-23
- Nature paper finds GPT-4 can predict social science experiment results with r=0.85 — JeremyNguyenPhD · 2026-07-24
4 near-duplicate retellings: RobbWiller · RobbWiller · RobbWiller · RobbWiller