CruxBench lands at NeurIPS: frontier LLMs barely beat random at asking the right questions

mengyer · x · 2026-10-03

CruxBench is accepted to NeurIPS 2026. Instead of grading answers against fixed labels, it scores LLM-generated "crux" questions by Value of Information: how much each question updates beliefs about real forecasting targets. Ground truth comes from future world events, making it contamination-resistant by construction, open-ended, and grounded in quantified belief shifts. Across 293 forecasting questions and 8 models, VOI correlates strongly with capability (r=0.90), yet even frontier LLMs only narrowly beat a random-timing baseline — information discovery remains a hard frontier problem.

Original post →

More from Research

Research channel →