EurekaBench: GPT-6 Astra nearly matches humans on prediction but lags on scientific insights

geoffwolfe · x · 2026-10-04

ReasonCore unveiled EurekaBench, a new benchmark testing whether AI experiments can uncover mechanisms that explain observations, complementing SciCode (scientific code) and CritPt (research-level physics). It spans 26 problems and 306 insight questions across six domains, scoring scientific constraints, predictive accuracy, and mechanistic insights separately.

Key finding: GPT-6 Astra nearly matches the human reference on prediction (47.4% vs 48.8%) but trails sharply on insights (29.4% vs 69.7%) — fitting data isn't understanding it. ReasonCore is building a dataset around the benchmark.

Related event: EurekaBench: AI Excels at Problem-Solving but Lacks Scientific Insight(2 posts)→

Original post →

More from Research

Research channel →