GPT-4 stays accurate on unpublished social science studies and open-weight models also perform well
RobbWiller · x · 2026-07-23
Model comparison
The paper reports that GPT-4’s predictions remained strongly correlated with observed effects even when the experiments were not public by the training cutoff or were still unpublished in June 2025.
Robustness
- GPT-4: raw correlation around .92 in the comparison plot
- High correlations held for experiments from multiple social-science domains
- Open-weight models also performed well, with correlations in the .81–.86 range in the figure
Takeaway
The result is not just a memorization artifact: performance stayed high on unpublished or later-published studies as well.
Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(18 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11