A 70-experiment benchmark finds GPT-4’s predicted effects at r=.85 versus reality
RobbWiller · x · 2026-07-23
Method and result
The authors evaluated GPT-4 on 70 preregistered experiments from TESS and replications, covering 469 effects and 119,330 participants.
They prompted the model with the study materials and demographic profiles, analyzed the simulated responses, and compared the inferred treatment effects with the real ones. The resulting estimates were highly correlated with the actual effects: r = .85, adjusted r = .92.
Broader pattern
Performance stayed strong across many social-science subfields, not just one narrow task.
Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(15 posts)→
More from Research
- Robotics paper says VLA and world models are not enough for grounded supervision — hbouammar · 2026-07-23
- AI could compress decades of biomedical research into days, says Derya Unutmaz — DeryaTR_ · 2026-07-23
- OpenAI and Apollo show models may optimize graders, not user intent — rohanpaul_ai · 2026-07-23
- AI Autonomously Disproves Decades-Old Math Conjectures: The Singularity's Opening Phase — imjustnewatai · 2026-07-23
- Applied Math Dominates AI, But Why Does Gradient Descent Actually Work? — fkasummer · 2026-07-23
- Cursor’s Composer 2.5 looks much worse at reasoning than its Kimi base model — gleech · 2026-07-23