LLM forecasts were about 2x too large and weaker on field experiments
RobbWiller · x · 2026-07-23
This reply highlights the main limitations of the LLM forecasts.
- On average, LLM predictions were about 2x too large.
- Accuracy was weaker for small effects, stable pre-existing attitudes, field experiments, and non-text treatments.
- The authors conclude that LLM forecasts can inform research, but should not replace human evidence, especially in harder-to-predict domains.
Related event: Nature Study: GPT-4 Can Predict Social Science Experiment Results(18 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11