SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis

SciConBench has been accepted by NeurIPS 2026, with the paper already on arXiv. The benchmark evaluates AI's ability to synthesize scientific conclusions in the health domain. Key findings: the best model's factual F1 on the scientific conclusion synthesis task is only 0.337; the top scorer, GPT-6.1 Sol, reaches just 41.1 on conclusion synthesis, with most models clustered around 0.30–0.33. Meanwhile, 61.1%–90% of AI-synthesized conclusions contain at least one fact contradicting expert-written Cochrane reviews. The team plans to run evaluation rounds every two months and is publicly seeking API access and funding support.

Confirmed

Why it matters

2026-10-07 ~ 2026-10-07 · 10 related posts

Primary sources