CruxBench lands at NeurIPS: frontier LLMs barely beat random at asking the right questions
mengyer · x · 2026-10-03
CruxBench is accepted to NeurIPS 2026. Instead of grading answers against fixed labels, it scores LLM-generated "crux" questions by Value of Information: how much each question updates beliefs about real forecasting targets. Ground truth comes from future world events, making it contamination-resistant by construction, open-ended, and grounded in quantified belief shifts. Across 293 forecasting questions and 8 models, VOI correlates strongly with capability (r=0.90), yet even frontier LLMs only narrowly beat a random-timing baseline — information discovery remains a hard frontier problem.
More from Research
- Causal Inference Assumptions: What to Check Before You Run the Model — Nasereliver · 2026-10-03
- V4.1 trains smoothly without dense warmup tuning: more complex algorithm, simpler information flow — teortaxesTex · 2026-10-03
- Teaching AI to Drive with Evolution: An Intuitive Neuroevolution Explainer — egehancry · 2026-10-03
- Why AI can't solve the mystery of time: training presupposes the very clock it must explain — johnseach · 2026-10-03
- arXiv trends suggest AI uplift hits quantum physics within a year, all physics by late 2029 — cephaloform · 2026-10-03
- Berkeley's Humanoid Intelligence Center wins both tracks at IROS 2026 RoCo Challenge — berkeley_ai · 2026-10-03