GPT-4 stays accurate on unpublished social science studies and open-weight models also perform well
RobbWiller · x · 2026-07-23
Model comparison
The paper reports that GPT-4’s predictions remained strongly correlated with observed effects even when the experiments were not public by the training cutoff or were still unpublished in June 2025.
Robustness
- GPT-4: raw correlation around .92 in the comparison plot
- High correlations held for experiments from multiple social-science domains
- Open-weight models also performed well, with correlations in the .81–.86 range in the figure
Takeaway
The result is not just a memorization artifact: performance stayed high on unpublished or later-published studies as well.
Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(15 posts)→
More from Research
- AI could compress decades of biomedical research into days, says Derya Unutmaz — DeryaTR_ · 2026-07-23
- OpenAI and Apollo show models may optimize graders, not user intent — rohanpaul_ai · 2026-07-23
- AI Autonomously Disproves Decades-Old Math Conjectures: The Singularity's Opening Phase — imjustnewatai · 2026-07-23
- Applied Math Dominates AI, But Why Does Gradient Descent Actually Work? — fkasummer · 2026-07-23
- Cursor’s Composer 2.5 looks much worse at reasoning than its Kimi base model — gleech · 2026-07-23
- Lanyon says its neurosymbolic solver is 20–250x faster than frontier models — burny_tech · 2026-07-23