Nature Study: GPT-4 Can Predict Social Science Experiment Results

A study published in *Nature* suggests that large language models can predict the outcomes of social science experiments. Researchers fed experimental materials and participant demographics into GPT-4, prompting the model to simulate participant responses. They then estimated treatment effects and compared them with actual results. This is significant because it shifts the utility of LLMs from mere "text generation" to "predicting research outcomes."

Key Findings

The authors evaluated the model on 70 pre-registered social science experiments, covering 469 effects and 119,330 participants using data from TESS and replication studies. According to Robb Willer, the effect sizes generated by GPT-4 showed a high correlation with real effects (0.85); after adjusting for uncertainty in effect sizes, the correlation reached 0.92.

Robustness and Open Materials

Robb Willer added that this performance isn't just "memorizing published results": GPT-4 maintained high predictive correlation even for studies unpublished before its training cutoff, or those still unpublished by June 2025. The authors also released the Nature paper, a free manuscript, an interactive demo, alongside the data and code, explicitly encouraging the community to further test where this method works, where it fails, and how to use it responsibly.

2026-07-23 ~ 2026-07-23 · 15 related posts

2 near-duplicate retellings: RobbWiller · RobbWiller