LLM forecasts were about 2x too large and weaker on field experiments
RobbWiller · x · 2026-07-23
This reply highlights the main limitations of the LLM forecasts.
- On average, LLM predictions were about 2x too large.
- Accuracy was weaker for small effects, stable pre-existing attitudes, field experiments, and non-text treatments.
- The authors conclude that LLM forecasts can inform research, but should not replace human evidence, especially in harder-to-predict domains.
Related event: Nature Study: LLMs Can Predict Social Science Experiment Results(14 posts)→
More from Research
- Science paper turns memristor drift into a 2.12 ms neural dynamical system — jiqizhixin · 2026-07-23
- W&B says world action models can show their “dream” before choosing actions — wandb · 2026-07-23
- W&B says NVIDIA’s 14B DreamZero model fails hard tasks despite clean scalar metrics — wandb · 2026-07-23
- Winning NVIDIA Challenge Solution: Memorize Patterns, Search and Verify at Inference — NVIDIAAI · 2026-07-23
- UW-Madison Receives Multiple DOE Genesis Awards for AI in Energy — KyleCranmer · 2026-07-23
- LMSYS's New OPD Training Makes Qwen3.5 Reason 3X Faster with Improved Accuracy — ying11231 · 2026-07-23