A 70-experiment benchmark finds GPT-4’s predicted effects at r=.85 versus reality
RobbWiller · x · 2026-07-23
Method and result
The authors evaluated GPT-4 on 70 preregistered experiments from TESS and replications, covering 469 effects and 119,330 participants.
They prompted the model with the study materials and demographic profiles, analyzed the simulated responses, and compared the inferred treatment effects with the real ones. The resulting estimates were highly correlated with the actual effects: r = .85, adjusted r = .92.
Broader pattern
Performance stayed strong across many social-science subfields, not just one narrow task.
Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(18 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11