Nature Study: GPT-4 Can Predict Social Science Experiment Results
A study published in *Nature* suggests that large language models can predict the outcomes of social science experiments. Researchers fed experimental materials and participant demographics into GPT-4, prompting the model to simulate participant responses. They then estimated treatment effects and compared them with actual results. This is significant because it shifts the utility of LLMs from mere "text generation" to "predicting research outcomes."
Key Findings
The authors evaluated the model on 70 pre-registered social science experiments, covering 469 effects and 119,330 participants using data from TESS and replication studies. According to Robb Willer, the effect sizes generated by GPT-4 showed a high correlation with real effects (0.85); after adjusting for uncertainty in effect sizes, the correlation reached 0.92.
Robustness and Open Materials
Robb Willer added that this performance isn't just "memorizing published results": GPT-4 maintained high predictive correlation even for studies unpublished before its training cutoff, or those still unpublished by June 2025. The authors also released the Nature paper, a free manuscript, an interactive demo, alongside the data and code, explicitly encouraging the community to further test where this method works, where it fails, and how to use it responsibly.
2026-07-23 ~ 2026-07-23 · 15 related posts
- GPT-4 Accurately Predicts Social Science Experiment Results, Nature Study Shows — RobbWiller · 2026-07-23
- Nature paper finds LLMs can predict social science experiments with r=.85 — RobbWiller · 2026-07-23
- A 70-experiment benchmark finds GPT-4’s predicted effects at r=.85 versus reality — RobbWiller · 2026-07-23
- GPT-4 stays accurate on unpublished social science studies and open-weight models also perform well — RobbWiller · 2026-07-23
- GPT-4 matches 2,659 human forecasters and improves when averaged with them — RobbWiller · 2026-07-23
- LLM forecasts were about 2x too large and weaker on field experiments — RobbWiller · 2026-07-23
- LLM-only pilots cost under $1 and rival ~230-person human pilots — RobbWiller · 2026-07-23
- 76% of surveyed social scientists would use AI for power analysis or pilot tests — RobbWiller · 2026-07-23
- Survey finds social scientists worry most about bias, reproducibility, and replacement — RobbWiller · 2026-07-23
- Researchers open a demo for forecasting social-science treatment effects — RobbWiller · 2026-07-23
- [source] Nature paper finds LLMs can predict social-science experiment results — RobbWiller · 2026-07-23
- Nature piece says LLMs can forecast social-science experiment outcomes — RobbWiller · 2026-07-23
- [source] Nature study packages paper, demo, data and code for predicting social science results — RobbWiller · 2026-07-23
2 near-duplicate retellings: RobbWiller · RobbWiller