SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis
SciConBench has been accepted by NeurIPS 2026, with the paper already on arXiv. The benchmark evaluates AI's ability to synthesize scientific conclusions in the health domain. Key findings: the best model's factual F1 on the scientific conclusion synthesis task is only 0.337; the top scorer, GPT-6.1 Sol, reaches just 41.1 on conclusion synthesis, with most models clustered around 0.30–0.33. Meanwhile, 61.1%–90% of AI-synthesized conclusions contain at least one fact contradicting expert-written Cochrane reviews. The team plans to run evaluation rounds every two months and is publicly seeking API access and funding support.
Confirmed
- SciConBench accepted at NeurIPS 2026, paper on arXiv, released by @manoelribeiro's team
- The benchmark contains 9.11K questions, with gold answers drawn from conclusions of expert-written systematic reviews (Cochrane)
- 61.1%–90% of AI-synthesized conclusions contain at least one fact contradicting Cochrane evidence
- Claude Opus 5.5 ranks among the top in the contradiction leaderboard
- Highest score on conclusion synthesis is GPT-6.1 Sol's 41.1 factual F1, with most models around 0.30–0.33
- Methodology: SciConHarness blocks access to the target ground-truth Cochrane reviews and other sources that could leak answers, forcing models to genuinely synthesize evidence across studies rather than retrieve answers directly
- Uses a fixed core set plus a monthly rolling panel of newly published reviews to distinguish genuine model progress from data leakage
- The team plans to run evaluations every two months over the next 1–2 years, tracking newly released frontier models, and the method can extend to scientific domains beyond health
- The team is publicly asking model providers for API access or evaluation credits/funding
Why it matters
- The authors note that most existing benchmarks fail to reflect how AI actually performs in high-stakes settings like health research; SciConBench focuses on "scientific conclusion synthesis"—the key task of identifying and weighing evidence across multiple studies to answer research questions
- Results show that conclusion synthesis in the medical domain remains a significantly unsolved challenge for frontier models, with limited progress
- The rolling panel mechanism offers a reusable evaluation paradigm for separating real capability gains from training-data leakage
2026-10-07 ~ 2026-10-07 · 10 related posts
Primary sources
- [source] SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07
- Most benchmarks miss how AI performs in high-stakes health research synthesis — manoelribeiro · 2026-10-07
- SciConHarness blocks ground-truth sources to force models to synthesize, not look up — manoelribeiro · 2026-10-07
- SciConHarness blocks answer sources to force genuine model synthesis — manoelribeiro · 2026-10-07
- SciConBench uses rolling review panels to separate real gains from leakage — manoelribeiro · 2026-10-07
- Frontier models barely improve at medical synthesis, best F1 just 41.1 — manoelribeiro · 2026-10-07
- 61–90% of AI-synthesized medical conclusions contain factual errors, SciConBench finds — manoelribeiro · 2026-10-07
- SciConBench to run bi-monthly, seeks API access and funding support — manoelribeiro · 2026-10-07
- [source] SciConBench: 61%-90% of AI-Synthesized Conclusions Contradict Cochrane Reviews — manoelribeiro · 2026-10-07
- [source] SciConBench Team to Rerun Evaluations Every Two Months, Seeks Funding — manoelribeiro · 2026-10-07